Skip to main content

Overview

Iris Control is a control layer for the LLM calls your own applications make. Your apps send requests to one OpenAI-compatible endpoint, and Iris Control applies your policy to every call before and after the model runs:
  • Routing and failover. Each use case has a primary model and fallbacks. If a provider errors or times out, the next model answers.
  • Data scoping. Retrieval is limited to the product or tenant a call belongs to, so an answer only uses the data it is allowed to see.
  • Tool limits. Each workflow can call only the tools you allow.
  • Grounding checks. Answers are checked against approved sources before they are returned. Unsupported answers are flagged or replaced with your handoff message.
  • Call log and evals. Every prompt, answer, model, and policy decision is logged, and you can score models against your own questions.
Try it before you connect anything: the Iris Control live demo runs model failover, data scoping, and grounding checks on live models, with no signup.

How it works

  1. Your application calls Iris Control with an OpenAI-compatible request. The model field names a use case (for example support-answers), not a specific model.
  2. Iris Control looks up the use case’s policy and applies it: scoped retrieval, allowed tools, and the model chain.
  3. The primary model answers. If its provider fails, Iris Control retries the call on the next model in the chain.
  4. The answer runs through the grounding check, when the use case has it turned on.
  5. Iris Control returns the answer in the standard OpenAI response format, with headers that show which model answered and what the policy did, and writes the call to the call log.

Quickstart

Step 1: Get an API key

API keys are provided by your Iris Control point of contact. Send them from your backend only, never from browser or mobile code.

Step 2: Add your model providers

In the Iris Control dashboard, open Providers and add the providers you use: OpenAI, Anthropic, Google Gemini, Azure OpenAI, AWS Bedrock, or a self-hosted model with an OpenAI-compatible endpoint. Provider keys are stored encrypted and are never returned by the API.

Step 3: Create a use case

In the dashboard, open Use cases and create one for each LLM feature you run, such as support-answers or invoice-agent. Set its model chain and, as needed, data scoping, tools, and grounding checks. The settings are described in Use case policies.

Step 4: Point your client at Iris Control

Change the base URL and API key in the OpenAI client your application already uses. Set model to the use case name.
Streaming (stream: true), tool calling, and structured outputs work the same way as with the OpenAI API.

Request headers

Tag each call so the policy can scope it and the call log can group it.

Response headers

Every response carries the outcome of the policy.

Use case policies

Each use case has one policy. You can edit it in the dashboard, or export and import it as YAML.

Models and failover

models.chain lists the models for the use case in order. Iris Control calls the first one. If the call fails with a provider error, a rate limit, or a timeout longer than attempt_timeout_seconds, it retries on the next model in the chain. Every attempt is recorded in the call log with its model, outcome, and latency, and X-Iris-Attempts shows how many models were tried. To change models, edit the chain. Your application code does not change, because it only names the use case.

Data scoping

When retrieval.enabled is on, Iris Control searches your connected knowledge before calling the model and adds the matching passages to the request. Knowledge comes from the data sources you connect, such as help centers, Confluence, websites, and uploaded content. retrieval.scope_by sets which request headers limit the search:
  • product: only knowledge tagged with the X-Iris-Product value is searched. Articles can belong to several products.
  • tenant: only knowledge tagged with the X-Iris-Tenant value is searched.
A call with a scope header that matches no knowledge gets no retrieved passages, so the model has nothing outside that scope to answer from. If you run your own retrieval, leave retrieval.enabled off and send the passages you retrieved with the request. Iris Control uses them for the grounding check:

Tool limits

tools.workflows maps each workflow to the tools it may call. For a request with an X-Iris-Workflow header:
  • Tools in the request that are not on the workflow’s list are removed before the model sees them.
  • A tool call in the model’s response to a tool that is not on the list is removed from the response.
Both cases are listed in X-Iris-Tools-Removed and recorded in the call log. A request without X-Iris-Workflow keeps the tools it sends.

Grounding checks

The grounding check runs after the model answers and before the response is returned. It runs two checks against the sources for the call, whether retrieved by Iris Control or sent in iris.sources:
  • Link check. Every link in the answer must point to one of the sources. Links that can’t be matched to a source are reported as unapproved.
  • Claim check. A judge model checks each claim in the answer against the source text and returns a verdict, a grounding score from 0 to 5, and any unsupported claims.
grounding.mode decides what happens when an answer fails either check:

Grounding check API

To check answers produced outside Iris Control, call the grounding check directly.
Response:
Send up to 5 sources per request.

Call log

Every call appears in the dashboard under Call log, with:
  • The use case, product, tenant, workflow, and your X-Iris-Trace-Id
  • Each model attempt, with its outcome and latency
  • The prompt, the retrieved or supplied sources, and the answer
  • Tools removed by policy
  • The grounding check result
Filter by use case, tenant, model, or grounding result, and export entries as CSV. Set how long prompts and answers are kept with logging.retention_days, or turn off logging.store_prompts to keep only metadata.

Evals on your own data

Compare models on your own questions before you change a use case’s model chain.
  1. In the dashboard, open Evals and upload a CSV with question and, optionally, expected_answer columns, or select calls from the call log.
  2. Pick the models to compare.
  3. Iris Control runs every question through each model with the use case’s retrieval and grounding settings, and scores each answer for grounding and, when an expected answer is given, correctness.
The scorecard shows accuracy, grounded rate, latency, and cost per model, so you can pick the model chain on your own data rather than a public benchmark.

Errors

Iris Control returns errors in the OpenAI error format. A blocked grounding check is not an error. It returns 200 with the handoff message and X-Iris-Grounding: blocked.