Routing and fallback
Resolve direct models, reusable routes, virtual route models, and direct fallbacks.
Every chat request resolves either a direct model or a Nirmos route.
Direct model
Use a provider-qualified identifier when model choice belongs in application code:
await nirmos.gateway.chat.create({
model: "openai/gpt-5.4-mini",
messages: [{ role: "user", content: "Hello" }],
});await client.chat.completions.create({
model: "openai/gpt-5.4-mini",
messages: [{ role: "user", content: "Hello" }],
});If an unqualified model exists for multiple providers, the gateway rejects it as ambiguous.
Nirmos route
Routes move target selection and fallback policy into Nirmos. The SDK can address a route directly; OpenAI-compatible clients use its virtual model alias.
await nirmos.gateway.chat.create({
route: "production-chat",
messages: [{ role: "user", content: "Hello" }],
});await client.chat.completions.create({
model: "nirmos/production-chat",
messages: [{ role: "user", content: "Hello" }],
});nirmos/production-chat is not a provider model. The gateway removes the prefix, looks up an enabled active route with slug production-chat, filters targets by required capabilities, and applies the route strategy.
Implemented route strategies
| Strategy | Behavior |
|---|---|
weighted | Selects among enabled targets using their positive configured weights |
round_robin | Advances across enabled targets for the current gateway process |
Targets that are unavailable or do not support required capabilities such as chat, vision, tools, or structured output are excluded.
Routing hints versus route strategies
The SDK's routingStrategy field accepts additional extensible values, but configured routes in the current gateway execute weighted or round_robin. Do not assume a hint creates a route or replaces its configured strategy.
Route fallback
An enabled route fallback policy contains ordered targets, trigger error types, a maximum attempt count, and a timeout. The gateway retries the selected target according to its execution policy, then advances through eligible fallback targets while the configured failure conditions allow it.
The route configuration—not the client request—defines those fallback targets.
Direct-model fallback
The Nirmos SDK exposes ordered fallbacks for a direct request:
await nirmos.gateway.chat.create({
model: "openai/gpt-5.4-mini",
fallbackModels: [
"google/gemini-3.5-flash",
"anthropic/claude-sonnet-4-6",
],
messages,
});Each fallback model must exist and support the capabilities required by the request. This is a Nirmos extension, so it is not shown as a standard OpenAI SDK option.
Resolution summary
openai/gpt-5.4-mini → direct provider-qualified model
nirmos/production-chat → active route with slug production-chat
route: production-chat → the same route through the Nirmos SDKThe Nirmos SDK validates that model and route are mutually exclusive before sending the request.