New: Prompt Evaluations now live

The infrastructure layer for every AI application.

Nirmos gives teams one platform to access any model, route requests, reduce costs, build agents, and observe every interaction.

Simulated traffic through a Nirmos route
128 requests15% from cache0 fallbacks
Your appmodel: "chat-default"
Nirmos gateway
  • Auth and limits
  • Cache lookup
  • Route and fallback
  • Trace
chat-default: weighted 50/30/20, then fall back

Select a provider to simulate an outage and watch Nirmos fall back.

Six most recent simulated requests. Times, providers, status, latency, tokens, and cost.
TimeModelStatusLatencyTokensCost
13:30:18OpenAI200617 ms662$0.0026
13:30:15Anthropic200654 ms639$0.0027
13:30:12Google200691 ms616$0.0028
13:30:09OpenAI200728 ms593$0.0029
13:30:06Cachecache hit8 ms570$0.0000
13:30:03Anthropic200802 ms547$0.0031

THE SPACE BETWEEN CODE AND MODELS

Build your product.
We’ll handle the in-between.

A model call is only the beginning. Nirmos brings the routing, prompt management, evaluations, and visibility around it into one infrastructure layer.

Your application
Nirmos infrastructure
Your AI providers

Every model call.
A little more under control.

From the first prompt to the production request, give your team the tools to build, test, and improve.

Test prompts before you ship.

Evaluate outputs against explicit criteria. Compare results and catch broken contracts before a prompt change reaches your users.

Explore prompt evaluations

One API. More possibilities.

Access providers through one gateway, with routing and ordered fallback to keep requests moving.

Explore the gateway

Prompts with a past. And a future.

Version your instructions, retrieve the active prompt, or pin a version for reproducible runs.

Meet managed prompts

Make repeat requests lighter.

Reuse cached responses for exact request matches and avoid another provider call.

See caching in action

See what happened.

Get request and trace IDs, provider details, latency, and usage metadata for your model calls.

Explore request metadata

MANAGED PROMPTS

Prompts you can version, test, and ship.

Write prompts once, fill them with typed variables, and publish a version when it is ready. Your application keeps calling the same name.

Versioned

Every edit is a version. Publish when it is ready, and roll back in one click.

Typed variables

Strings, numbers, booleans, enums, and arrays, validated before a request goes out.

Organized

Keep prompts in folders and reuse the same prompt across apps and environments.

Prompts / support / triage-reply
v13 draftv12 published
Prompt template
You are a support specialist for Acme Cloud.

Write a {{tone}} reply to {{customer_name}} about their open ticket.
Keep it under {{max_words}} words.
Include next steps: {{include_steps}}
Variables
Rendered prompt, as sent to the model
You are a support specialist for Acme Cloud.

Write a friendly reply to Priya about their open ticket.
Keep it under 80 words.
Include next steps: true

v13 draft. v12 published.

PROMPT EVALUATIONS

Know which prompt and model is better before your users do.

Run prompts and models against a dataset, score every output with the evaluators you configure, and compare candidates side by side.

Evaluations / triage-reply on support-tickets
30 cases
LLM judge: helpfulness of 4 or moreMentions the refund windowUnder 120 wordsLatency under 2 s
Prompt v12 on OpenAIBaseline
—% pass
Median latency 1.41 sCost per 1k $4.10
Prompt v13 on AnthropicCandidate
—% pass
Evaluating candidate…
Median latency 1.22 sCost per 1k $5.30
Prompt v13 on GoogleCandidate
—% pass
Evaluating candidate…
Median latency 0.96 sCost per 1k $2.40
Passed every evaluatorFailed at least oneEach square is one test case. Sample results.

Evaluation ready.

One API for every model.

A familiar API. A new level of control.

Keep your OpenAI client. Point it at Nirmos, use your Nirmos API key, and choose a provider-qualified model. Use the Nirmos SDK when you need managed prompts and normalized metadata.

OpenAI-compatible chat completions
Provider-qualified models
Weighted and round-robin routes
Ordered provider fallback
Streaming responses
Managed prompts with Nirmos SDK
Make your first request
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.nirmos.com/v1",
  apiKey: process.env.NIRMOS_API_KEY!,
});

const completion = await client.chat.completions.create({
  model: "openai/gpt-5.4-mini",
  messages: [
    { role: "user", content: "Explain an AI gateway." },
  ],
  max_tokens: 200,
});
Same OpenAI client. Nirmos endpoint. Your API key.

Put one layer between your app and every model.

Change a base URL, and start routing, testing, and tracing your AI traffic today.