Nirmos

Routing and fallback

Resolve direct models, reusable routes, virtual route models, and direct fallbacks.

Every chat request resolves either a direct model or a Nirmos route.

Direct model

Use a provider-qualified identifier when model choice belongs in application code:

await nirmos.gateway.chat.create({
  model: "openai/gpt-5.4-mini",
  messages: [{ role: "user", content: "Hello" }],
});
await client.chat.completions.create({
  model: "openai/gpt-5.4-mini",
  messages: [{ role: "user", content: "Hello" }],
});

If an unqualified model exists for multiple providers, the gateway rejects it as ambiguous.

Nirmos route

Routes move target selection and fallback policy into Nirmos. The SDK can address a route directly; OpenAI-compatible clients use its virtual model alias.

await nirmos.gateway.chat.create({
  route: "production-chat",
  messages: [{ role: "user", content: "Hello" }],
});
await client.chat.completions.create({
  model: "nirmos/production-chat",
  messages: [{ role: "user", content: "Hello" }],
});

nirmos/production-chat is not a provider model. The gateway removes the prefix, looks up an enabled active route with slug production-chat, filters targets by required capabilities, and applies the route strategy.

Implemented route strategies

StrategyBehavior
weightedSelects among enabled targets using their positive configured weights
round_robinAdvances across enabled targets for the current gateway process

Targets that are unavailable or do not support required capabilities such as chat, vision, tools, or structured output are excluded.

Routing hints versus route strategies

The SDK's routingStrategy field accepts additional extensible values, but configured routes in the current gateway execute weighted or round_robin. Do not assume a hint creates a route or replaces its configured strategy.

Route fallback

An enabled route fallback policy contains ordered targets, trigger error types, a maximum attempt count, and a timeout. The gateway retries the selected target according to its execution policy, then advances through eligible fallback targets while the configured failure conditions allow it.

The route configuration—not the client request—defines those fallback targets.

Direct-model fallback

The Nirmos SDK exposes ordered fallbacks for a direct request:

await nirmos.gateway.chat.create({
  model: "openai/gpt-5.4-mini",
  fallbackModels: [
    "google/gemini-3.5-flash",
    "anthropic/claude-sonnet-4-6",
  ],
  messages,
});

Each fallback model must exist and support the capabilities required by the request. This is a Nirmos extension, so it is not shown as a standard OpenAI SDK option.

Resolution summary

openai/gpt-5.4-mini     → direct provider-qualified model
nirmos/production-chat  → active route with slug production-chat
route: production-chat  → the same route through the Nirmos SDK

The Nirmos SDK validates that model and route are mutually exclusive before sending the request.

On this page