The infrastructure layer for every AI application.
Nirmos gives teams one platform to access any model, route requests, reduce costs, build agents, and observe every interaction.
model: "chat-default"- Auth and limits
- Cache lookup
- Route and fallback
- Trace
chat-default: weighted 50/30/20, then fall backSelect a provider to simulate an outage and watch Nirmos fall back.
| Time | Model | Status | Latency | Tokens | Cost |
|---|---|---|---|---|---|
| 13:30:18 | OpenAI | 200 | 617 ms | 662 | $0.0026 |
| 13:30:15 | Anthropic | 200 | 654 ms | 639 | $0.0027 |
| 13:30:12 | 200 | 691 ms | 616 | $0.0028 | |
| 13:30:09 | OpenAI | 200 | 728 ms | 593 | $0.0029 |
| 13:30:06 | Cache | cache hit | 8 ms | 570 | $0.0000 |
| 13:30:03 | Anthropic | 200 | 802 ms | 547 | $0.0031 |
THE SPACE BETWEEN CODE AND MODELS
Build your product.
We’ll handle the in-between.
A model call is only the beginning. Nirmos brings the routing, prompt management, evaluations, and visibility around it into one infrastructure layer.
THE NIRMOS TOOLKIT
Every model call.
A little more under control.
From the first prompt to the production request, give your team the tools to build, test, and improve.
Test prompts before you ship.
Evaluate outputs against explicit criteria. Compare results and catch broken contracts before a prompt change reaches your users.
Explore prompt evaluationsOne API. More possibilities.
Access providers through one gateway, with routing and ordered fallback to keep requests moving.
Explore the gatewayPrompts with a past. And a future.
Version your instructions, retrieve the active prompt, or pin a version for reproducible runs.
Meet managed promptsMake repeat requests lighter.
Reuse cached responses for exact request matches and avoid another provider call.
See caching in actionSee what happened.
Get request and trace IDs, provider details, latency, and usage metadata for your model calls.
Explore request metadataMANAGED PROMPTS
Prompts you can version, test, and ship.
Write prompts once, fill them with typed variables, and publish a version when it is ready. Your application keeps calling the same name.
Versioned
Every edit is a version. Publish when it is ready, and roll back in one click.
Typed variables
Strings, numbers, booleans, enums, and arrays, validated before a request goes out.
Organized
Keep prompts in folders and reuse the same prompt across apps and environments.
You are a support specialist for Acme Cloud. Write a {{tone}} reply to {{customer_name}} about their open ticket. Keep it under {{max_words}} words. Include next steps: {{include_steps}}
You are a support specialist for Acme Cloud. Write a friendly reply to Priya about their open ticket. Keep it under 80 words. Include next steps: true
v13 draft. v12 published.
PROMPT EVALUATIONS
Know which prompt and model is better before your users do.
Run prompts and models against a dataset, score every output with the evaluators you configure, and compare candidates side by side.
Evaluation ready.
AI Gateway · Developer experience
One API for every model.
A familiar API. A new level of control.
Keep your OpenAI client. Point it at Nirmos, use your Nirmos API key, and choose a provider-qualified model. Use the Nirmos SDK when you need managed prompts and normalized metadata.
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.nirmos.com/v1",
apiKey: process.env.NIRMOS_API_KEY!,
});
const completion = await client.chat.completions.create({
model: "openai/gpt-5.4-mini",
messages: [
{ role: "user", content: "Explain an AI gateway." },
],
max_tokens: 200,
});
Put one layer between your app and every model.
Change a base URL, and start routing, testing, and tracing your AI traffic today.