How to use it
Call the LLM Gateway with an OpenAI-compatible client and choose a provider model or gateway alias.
Call the gateway
Point an OpenAI-compatible client at LLM_GATEWAY_ENDPOINT, authenticate with LLM_GATEWAY_API_KEY, and pass a gateway alias or a provider/model string as model. Use the Responses API for new work.
import asyncio
import os
from openai import AsyncOpenAI
client = AsyncOpenAI(
api_key=os.environ["LLM_GATEWAY_API_KEY"],
base_url=os.environ["LLM_GATEWAY_ENDPOINT"],
)
async def main() -> None:
baseline = await client.responses.create(
model="frontier-production",
input="Review this migration plan for hidden deployment risk.",
)
print(baseline.output_text)
pinned = await client.responses.create(
model="anthropic/claude-opus-5",
input="Review this changelog for user-facing risk.",
)
print(pinned.output_text)
routed = await client.responses.create(
model="fast-production",
input="Summarize the latest support ticket.",
)
print(routed.output_text)
if __name__ == "__main__":
asyncio.run(main())Keep provider credentials out of the project. The gateway variables are the only credentials your code needs.
LLM_GATEWAY_ENDPOINTis the unversioned gateway host. The canonical Responses route is/v1/responses; the gateway also mounts/responses, so a client with the host asbase_urlcan also reach/chat/completionsfor legacy code.- The gateway rejects request bodies that set
api_key,api_base, orbase_url, or the gateway-owned retry and metadata fields, with a 400. Keep those settings in client setup. - If a streamed response fails after streaming starts, the gateway sends a Responses
errorevent (upstream_errororstream_interrupted) and then ends the stream. An HTTP 200 alone does not mean the stream completed.
Choose a model
Aliases keep application code stable while Ciridae updates the models behind them. Start a new agent task on frontier-production, prove quality on representative evals, then test cheaper aliases against the same evals. Estimate total cost before scaling a batch.
| Model value | Use it for | Fallback on provider failure |
|---|---|---|
frontier-production | Establishing whether a new task is solvable | No; errors surface after retries |
strong-production | Proven tasks that need strong reasoning at higher volume | Yes, across providers |
fast-production | Proven chat, summaries, and short transformations | Yes, across providers |
tiny-production | Trivial classification and high-volume, low-risk work | Yes, across providers |
openai/..., anthropic/..., gemini/... | Behavior that needs one specific provider model | No |
Each alias has a primary model and an ordered fallback list defined in the gateway. On the Responses routes, the fallback aliases retry a retryable failure up to five times before moving to the next model. When a request omits reasoning effort, frontier-production and fast-production default to medium and strong-production and tiny-production default to low; an explicit effort wins. Provider/model strings go straight to that provider with no gateway reasoning default.
Use the OpenAI Agents SDK
The gateway supports the Responses API calls the OpenAI Agents SDK makes. Wrap a gateway-backed client in OpenAIResponsesModel, then change the model value without rewriting the agent.
import asyncio
import os
from agents import Agent, OpenAIResponsesModel, RunConfig, Runner
from openai import AsyncOpenAI
gateway_client = AsyncOpenAI(
api_key=os.environ["LLM_GATEWAY_API_KEY"],
base_url=os.environ["LLM_GATEWAY_ENDPOINT"],
)
agent = Agent(
name="Support triage",
instructions="Classify the request and recommend the next step.",
model=OpenAIResponsesModel(
model="strong-production",
openai_client=gateway_client,
),
)
async def main() -> None:
result = await Runner.run(
agent,
"Plan the next action for this workflow.",
run_config=RunConfig(tracing_disabled=True),
)
print(result.final_output)
if __name__ == "__main__":
asyncio.run(main())