Using OpenRouter
Reach hundreds of models from every major vendor through one OpenRouter API key, with gateway spend that matches your OpenRouter invoice.
OpenRouter is a marketplace rather than a single vendor: one key reaches models from OpenAI, Anthropic, Google, Meta, DeepSeek, Qwen, xAI and many more, and OpenRouter chooses an upstream host for each request. Adding it as a provider serves those models to your apps through the Universal API, on your OpenRouter account and your OpenRouter billing.
Three things set this provider apart, and each one changes something you'll type:
- Model references carry the vendor prefix, so they have two slashes.
- Claude models serve the Messages API, not the Responses API.
- Reported spend matches your OpenRouter invoice, because OpenRouter prices each request and the gateway bills from that figure.
Model references include the vendor prefix
Every OpenRouter model ID is vendor-qualified — anthropic/claude-sonnet-5,
openai/gpt-6-sol, meta-llama/llama-3.3-70b-instruct. A gateway model
reference prefixes your provider name onto that, so it has two slashes:
Code
The first segment is your provider name, and everything after the first slash is the model ID exactly as OpenRouter publishes it. This is the same shape Vertex AI uses.
Supported endpoints
Which endpoints a model serves depends on what kind of model it is. Claude is served on a different API from every other chat model, and embedding models serve only embeddings:
| Endpoint | Claude models (anthropic/…) | Other chat models | Embedding models |
|---|---|---|---|
/v1/chat/completions | ✅ (translated) | ✅ | ❌ |
/v1/messages | ✅ | ❌ | ❌ |
/v1/messages/count_tokens | ❌ | ❌ | ❌ |
/v1/responses | ❌ | ✅ | ❌ |
/v1/embeddings | ❌ | ❌ | ✅ |
The Responses API is stateless on OpenRouter
OpenRouter doesn't store responses. You can create one, but you can't retrieve
or delete it or list its input items, and OpenRouter rejects requests that set
store: true or previous_response_id. Send the full conversation with each
request instead.
OpenRouter doesn't offer token counting either, so /v1/messages/count_tokens
returns a 400 for every OpenRouter model.
Claude models don't serve the Responses API here
OpenRouter's own Responses endpoint will accept a Claude model. The gateway
doesn't allow it, so that every provider follows the same rule: /v1/messages
means Claude, and /v1/responses means everything else. To call a Claude model
from an OpenAI-compatible client, use /v1/chat/completions: the gateway
translates the request onto the Messages API for you.
Before you begin
You need an OpenRouter API key. Create one at
openrouter.ai/keys; it starts with sk-or-v1-.
You also need credit on the OpenRouter account the key belongs to, or requests come back with OpenRouter's own insufficient-credit error.
Add the provider
Adding or editing providers requires the Edit permission, granted to Zuplo account and project Admins—see Managing Providers.
-
Open Settings → AI Providers in your AI Gateway project in the Zuplo Portal.
-
Select Add Provider, then choose OpenRouter.
-
Give the provider a Provider Name. This is the first segment of every model reference, so a provider named
openrouterservesopenrouter/openai/gpt-6-sol. The name is permanent after creation. -
Paste your API Key. There is no endpoint or region to enter — OpenRouter is a single global host.
-
Select the models this provider should serve, then Save.
Routing shortcuts
OpenRouter lets you append a routing preference to any model ID:
| Suffix | Effect |
|---|---|
:nitro | Prefer the fastest upstream host |
:floor | Prefer the cheapest upstream host |
:online | Add web search to the request |
:exacto | Prefer hosts that call tools most accurately |
These work through the gateway and are forwarded to OpenRouter unchanged. They
don't appear in the model picker, because they aren't separate models — the
gateway resolves openrouter/openai/gpt-6-sol:nitro against the
openai/gpt-6-sol catalog entry and sends the suffix upstream.
Two consequences worth knowing:
- A model-filtering allow list matches literally. Allowing
openrouter/openai/gpt-6-soldoes not allowopenrouter/openai/gpt-6-sol:nitro. Add the suffixed reference explicitly if you want it. - Other suffixes aren't supported. OpenRouter's
:freevariants are separate models, and the gateway's OpenRouter catalog doesn't include them, so a request for one is rejected like any model outside your selections.
Call the models
Point any OpenAI-compatible client at your app URL with /v1 appended:
Code
Provider routing preferences
OpenRouter's provider option is forwarded, so you can express upstream
preferences per request:
Code
provider is forwarded on embeddings and the Responses API too. On chat
completions, transforms, reasoning and plugins are forwarded the same way.
The Responses API forwards plugins, and reasoning is a standard Responses
parameter.
The models fallback array is never forwarded
OpenRouter's models array (and route: "fallback") asks OpenRouter to
substitute a different model when the first one is unavailable. On the Messages
API the same array is called fallbacks. On chat completions and the Messages
API, the gateway rejects these parameters with a 400 rather than dropping
them silently. On embeddings and the Responses API, it drops them before the
request reaches OpenRouter.
They would route around the gateway itself: your model filtering wouldn't see the substitute, your budgets and per-model costs would be attributed to a model that never ran, and it would collide with the gateway's own fallback handling. Configure a backup model on the route instead.
Call Claude models on the Messages API
Claude models on OpenRouter serve Anthropic's native Messages API, so the
Anthropic SDK works against them directly. Set baseURL without /v1 — the
Anthropic SDK appends that itself — and pass your key as authToken:
Code
authToken, not apiKey: the SDK's apiKey sends an x-api-key header, which
the gateway doesn't read.
For the Vercel AI SDK, see the OpenRouter section of the AI SDK guide.
Costs
OpenRouter prices every request itself and reports the figure in the response, and the gateway records that figure rather than computing one from a stored rate card. Your reported spend therefore reconciles with your OpenRouter invoice.
This matters more here than for a single-vendor provider. A :floor request and
a :nitro request for the same model can run on different upstream hosts at
different prices, so there is no single per-model rate that would be correct for
both.
Responses carry an X-Cost-USD header alongside X-Cost-Source: provider,
which tells you the figure came from OpenRouter rather than from a rate card.
Troubleshooting
400 saying /v1/responses is not supported — you sent a Claude model to
/v1/responses. Use /v1/messages, or /v1/chat/completions, which the
gateway translates onto the Messages API; see
the table above.
400 naming /v1/messages — you sent a model other than Claude to
/v1/messages, which serves only Claude models here.
An error saying the model is not included in your model selections — the model isn't selected on this provider. Open the provider on the AI Providers page and enable it.
400 rejecting models, route or fallbacks — see the caution above;
use a route backup model instead.
400 from /v1/messages/count_tokens — OpenRouter doesn't offer token
counting.
404 from OpenRouter — usually a model ID typo. The ID after your provider
name must match OpenRouter's exactly, vendor prefix and all.
A key that won't save — OpenRouter keys start with sk-or-v1- and contain
no whitespace. Re-copy it from openrouter.ai/keys.
Next steps
- Universal API — the endpoints every provider shares
- Managing providers — editing models and keys
- AI Providers — the full provider and capability matrix