# How to use it (/llm-gateway/how-to-use-it)



## Call the gateway [#call-the-gateway]

Point an OpenAI-compatible client at `LLM_GATEWAY_ENDPOINT`, authenticate with `LLM_GATEWAY_API_KEY`, and pass a gateway alias or a provider/model string as `model`. Use the Responses API for new work.

```python title="gateway_client.py"
import asyncio
import os

from openai import AsyncOpenAI

client = AsyncOpenAI(
    api_key=os.environ["LLM_GATEWAY_API_KEY"],
    base_url=os.environ["LLM_GATEWAY_ENDPOINT"],
)


async def main() -> None:
    baseline = await client.responses.create(
        model="frontier-production",
        input="Review this migration plan for hidden deployment risk.",
    )
    print(baseline.output_text)

    pinned = await client.responses.create(
        model="anthropic/claude-opus-5",
        input="Review this changelog for user-facing risk.",
    )
    print(pinned.output_text)

    routed = await client.responses.create(
        model="fast-production",
        input="Summarize the latest support ticket.",
    )
    print(routed.output_text)


if __name__ == "__main__":
    asyncio.run(main())
```

Keep provider credentials out of the project. The gateway variables are the only credentials your code needs.

<Accordions>
  <Accordion title="Under the hood: routes and request rules">
    * `LLM_GATEWAY_ENDPOINT` is the unversioned gateway host. The canonical Responses route is `/v1/responses`; the gateway also mounts `/responses`, so a client with the host as `base_url` can also reach `/chat/completions` for legacy code.
    * The gateway rejects request bodies that set `api_key`, `api_base`, or `base_url`, or the gateway-owned retry and metadata fields, with a 400. Keep those settings in client setup.
    * If a streamed response fails after streaming starts, the gateway sends a Responses `error` event (`upstream_error` or `stream_interrupted`) and then ends the stream. An HTTP 200 alone does not mean the stream completed.
  </Accordion>
</Accordions>

## Choose a model [#choose-a-model]

Aliases keep application code stable while Ciridae updates the models behind them. Start a new agent task on `frontier-production`, prove quality on representative evals, then test cheaper aliases against the same evals. Estimate total cost before scaling a batch.

| Model value                                 | Use it for                                               | Fallback on provider failure     |
| ------------------------------------------- | -------------------------------------------------------- | -------------------------------- |
| `frontier-production`                       | Establishing whether a new task is solvable              | No; errors surface after retries |
| `strong-production`                         | Proven tasks that need strong reasoning at higher volume | Yes, across providers            |
| `fast-production`                           | Proven chat, summaries, and short transformations        | Yes, across providers            |
| `tiny-production`                           | Trivial classification and high-volume, low-risk work    | Yes, across providers            |
| `openai/...`, `anthropic/...`, `gemini/...` | Behavior that needs one specific provider model          | No                               |

<Accordions>
  <Accordion title="Under the hood: alias routing">
    Each alias has a primary model and an ordered fallback list defined in the gateway. On the Responses routes, the fallback aliases retry a retryable failure up to five times before moving to the next model. When a request omits reasoning effort, `frontier-production` and `fast-production` default to `medium` and `strong-production` and `tiny-production` default to `low`; an explicit effort wins. Provider/model strings go straight to that provider with no gateway reasoning default.
  </Accordion>
</Accordions>

## Use the OpenAI Agents SDK [#use-the-openai-agents-sdk]

The gateway supports the Responses API calls the OpenAI Agents SDK makes. Wrap a gateway-backed client in `OpenAIResponsesModel`, then change the model value without rewriting the agent.

```python title="gateway_agent.py"
import asyncio
import os

from agents import Agent, OpenAIResponsesModel, RunConfig, Runner
from openai import AsyncOpenAI

gateway_client = AsyncOpenAI(
    api_key=os.environ["LLM_GATEWAY_API_KEY"],
    base_url=os.environ["LLM_GATEWAY_ENDPOINT"],
)

agent = Agent(
    name="Support triage",
    instructions="Classify the request and recommend the next step.",
    model=OpenAIResponsesModel(
        model="strong-production",
        openai_client=gateway_client,
    ),
)


async def main() -> None:
    result = await Runner.run(
        agent,
        "Plan the next action for this workflow.",
        run_config=RunConfig(tracing_disabled=True),
    )
    print(result.final_output)


if __name__ == "__main__":
    asyncio.run(main())
```
