Nirmos

What is Nirmos

A practical overview of the Nirmos gateway, SDK, routing, observability, and managed prompts.

Nirmos sits between your application and model providers. Your application sends one request to Nirmos; the gateway authenticates it, resolves a model or route, calls the selected provider, and returns an OpenAI-compatible or SDK-normalized response.

Your application → Nirmos Gateway → OpenAI, Anthropic, Google, or a compatible provider
                         ├─ routing and fallback
                         ├─ request traces and usage
                         └─ response caching

What you use

SurfaceUse it when
Nirmos SDKYou want typed camelCase requests, normalized responses, streaming helpers, managed prompts, and Nirmos error classes.
OpenAI-compatible APIYou already use the OpenAI JavaScript client or another OpenAI-compatible tool and need chat, streaming, models, or embeddings.
Managed promptsPrompt content should be versioned and released independently of application code. This is a Nirmos SDK feature.

Gateway model access

Direct requests use a provider-qualified model such as openai/gpt-5.4-mini. This avoids ambiguity when more than one provider exposes a model with the same identifier.

Routes move provider and model selection into Nirmos. Call a route with the Nirmos SDK's route field, or use its virtual model name from any OpenAI-compatible client:

// Nirmos SDK
{ route: "production-chat", messages }

// OpenAI-compatible client
{ model: "nirmos/production-chat", messages }

An active route can select targets using the configured weighted or round-robin strategy. Its fallback policy can move a failed request to another enabled target.

Observability

Gateway responses include request and trace identifiers plus the selected provider, model, latency, usage, and cache metadata when available. The Nirmos SDK normalizes these fields on completion.metadata and completion.usage.

Store the request ID with your application logs. It is the join key between an application failure and the corresponding Nirmos trace.

Managed prompts

Managed prompts are not part of the OpenAI API. The Nirmos SDK can retrieve an active or pinned prompt version, render {{variables}} locally, cache reads in memory, and insert the rendered messages into a gateway request.

Server-side credentials

Both clients use a Nirmos API key. Keep it in trusted server code; never expose it in browser bundles or public environment variables.

Where to go next

On this page