AI Providers
Zuplo's AI Gateway supports integration with various AI providers, allowing you to leverage different models and services for your apps.
Supported Providers
Zuplo currently supports the following AI providers:
- OpenAI
- Anthropic
- Mistral
- xAI (Grok)
- Azure AI—Azure OpenAI and Microsoft Foundry resources, serving Claude and every Foundry model family from your own Azure subscription
- Bedrock Mantle—Amazon Bedrock's compatible-APIs endpoint, serving Claude models and models from many other vendors
- Bedrock Runtime—Amazon Bedrock's native API, for code that already calls Bedrock with an AWS SDK
- Vertex AI—Google Cloud's managed model platform, serving Gemini and Model Garden models on your own Google Cloud project
- OpenRouter—hundreds of models from every major vendor behind a single API key
- Jev—TypeSafe's System One model, which answers typed questions about text instead of generating it
- Zuplo Demo—a free, keyless provider for trying the gateway
- OpenAI-compatible Custom Providers (such as Qwen, Kimi, etc)
The following capabilities are supported across providers:
| Provider | Chat Completions | Embeddings | Responses | Messages |
|---|---|---|---|---|
| OpenAI | ✅ | ✅ | ✅ | ❌ |
| Anthropic | ✅ | ❌ | ❌ | ✅ |
| ✅ | ✅ | ❌ | ❌ | |
| Mistral | ✅ | ✅ | ❌ | ❌ |
| xAI | ✅ | ✅ | ❌ | ❌ |
| Azure AI | ✅ | ✅ | ✅ | ✅ |
| Bedrock Mantle | ✅ | ❌ | ✅ | ✅ |
| Bedrock Runtime | ❌ | ❌ | ❌ | ❌ |
| Vertex AI | ✅ | ✅ | ❌ | ✅ |
| OpenRouter | ✅ | ✅ | ✅ | ✅ |
| Jev | ❌ | ❌ | ❌ | ❌ |
| Zuplo Demo | ✅ | ❌ | ❌ | ❌ |
| OpenAI-compatible (Custom) | ✅ | ✅ | ❌ | ❌ |
Responses is the OpenAI Responses API (/v1/responses), and Messages is
the native Anthropic Messages API (/v1/messages). See the
Universal API for the endpoint each
capability maps to.
Bedrock Mantle's capabilities depend on the model family: Claude models serve Messages (plus chat completions through translation), while its other models serve chat completions and—per model—Responses. See Using Bedrock Mantle.
Bedrock Runtime and Jev serve none of these capabilities, and the table above says so. Bedrock Runtime serves Amazon Bedrock's native API instead, at its own paths, so your existing AWS SDK code keeps working with the gateway in front of it. See Using Bedrock Runtime for the operations it serves, and add Bedrock Mantle alongside it if you also want Bedrock models on the Universal API.
Jev serves TypeSafe's System One API instead, at /v1/systemone. You send it
text and typed questions, and it returns typed answers rather than generated
text. See Using Jev.
Azure AI's capabilities split by model family, and within the OpenAI-compatible family they split per deployed model: a chat model serves chat completions and— per model and version—Responses, while an embedding model serves embeddings and neither of the others. Claude serves Messages (plus chat completions through translation) and needs a Microsoft Foundry resource, since Claude can't be deployed on an Azure OpenAI resource. Azure also addresses models by deployment name rather than by published model ID, so a deployment named differently from the model it serves needs mapping for the gateway to price it. See Using Azure AI.
A custom provider must serve chat completions and embeddings under a /v1 path
segment on its API URL, and you enter that URL as an origin root—without the
/v1 suffix vendors usually publish. See
Custom Providers for the exact contract.
Apps reference a provider's models as providerName/model—for example
openai/gpt-6-luna or anthropic/claude-sonnet-5—where providerName is the
name you give the provider configuration. See the
Universal API.
Azure AI
Azure AI serves models from your own Azure resource. One provider configuration covers both kinds of resource, and which capabilities apply depends on the resource and the model:
- An Azure OpenAI resource deploys OpenAI models and serves chat completions, embeddings, and the Responses API.
- A Microsoft Foundry resource serves those plus every Foundry model family—Grok, DeepSeek, Llama, Mistral, Phi, Kimi and more—and serves Claude models on the native Anthropic Messages API, with chat completions through the gateway's translation.
Two things make Azure different from every other provider:
- It addresses deployments, not published model IDs. You name deployments
when you create them in Azure, and that name is what goes in the
modelfield. A deployment whose name differs from the model it serves needs a mapping, or the gateway can't price its usage. - Its endpoint host embeds your resource name. The provider dialog asks for an Azure Resource Name and a host family rather than a full URL, and it takes a resource key—not a Microsoft Entra ID token.
For prerequisites, setup steps, code examples, and troubleshooting, see Using Azure AI.
Bedrock Mantle
Bedrock Mantle is Amazon Bedrock's compatible-APIs endpoint. One regional endpoint and one long-term Bedrock API key serve two model families, and which capabilities apply depends on the model you call:
- Claude models (for example
anthropic.claude-sonnet-5) serve the native Anthropic Messages API, and chat completions through the gateway's translation. - OpenAI-compatible models—everything else Mantle serves, from open-weight
models such as
openai.gpt-oss-120bto frontier models such asopenai.gpt-5.6-sol—serve chat completions and, per model, the OpenAI Responses API.
Mantle endpoints are regional, so the provider dialog asks for an AWS Region
instead of an endpoint URL, and it accepts only long-term Bedrock API keys,
which start with ABSK. For prerequisites, setup steps, code examples, and
troubleshooting, see Using Bedrock Mantle.
Bedrock Runtime
Bedrock Runtime is Amazon Bedrock's native API—the one the AWS SDKs call.
Unlike every other provider, it doesn't serve the
Universal API: it serves Bedrock's own operations
(Converse, ConverseStream, InvokeModel, and
InvokeModelWithResponseStream) at Bedrock's own paths.
That makes it the provider to choose when you already have code calling Bedrock with an AWS SDK. You point the client's endpoint at your app and give it the app's API key; every call site stays as it is, and the gateway adds authentication, model filtering, and usage metering in front of your own AWS account.
It authenticates to AWS with a long-term IAM access key pair rather than a Bedrock API key, so it needs its own provider configuration even if you also use Bedrock Mantle on the same account. For prerequisites, setup steps, code examples, and troubleshooting, see Using Bedrock Runtime.
Vertex AI
Vertex AI is Google Cloud's managed model platform. One provider configuration serves both the Gemini family and Vertex's Model Garden partner models—DeepSeek, Qwen, GLM, Kimi, MiniMax, Gemma, and GPT-OSS—from your own Google Cloud project, on your Google Cloud billing.
Two things make its setup different from every other provider:
- It authenticates with a service account, not an API key. Vertex's prediction endpoints reject API keys, so the dialog asks for a service account JSON key file and the gateway exchanges it for short-lived access tokens.
- It needs a Google Cloud project as well as a location. Vertex endpoints are per-location, and the project isn't part of the endpoint, so the dialog asks for both.
Every Vertex model ID is publisher-qualified, so a model reference has two
slashes: a provider named vertexai serves vertexai/google/gemini-3.7-flash,
vertexai/qwen/qwen3-coder-480b-a35b-instruct-maas, and
vertexai/anthropic/claude-haiku-4-5@20251001. Model availability varies by
location—most Model Garden models are served only from the global location,
and Google's regional endpoints serve Claude Sonnet 4.6 and earlier only, so
global is the location that serves every Claude model. Vertex's Claude models
serve the native Anthropic Messages API (/v1/messages) and nothing else: they
aren't available on /v1/chat/completions, and each one must be enabled in
Model Garden for your Google Cloud project before it answers.
For prerequisites, Google Cloud setup, provider steps, code examples, and troubleshooting, see Using Vertex AI.
OpenRouter
OpenRouter is a marketplace, not a single vendor: one API key reaches hundreds of models from OpenAI, Anthropic, Google, Meta, DeepSeek, Qwen, xAI and many more, and OpenRouter picks an upstream host for each request.
Three things are specific to it:
- Model references have two slashes. Every OpenRouter model ID is
vendor-qualified, so a provider named
openrouterservesopenrouter/anthropic/claude-sonnet-5andopenrouter/openai/gpt-6-sol. This is the same shape Vertex AI uses. - Claude models are on the Messages API. Models whose ID begins
anthropic/serve the native Anthropic Messages API (/v1/messages) and a translated/v1/chat/completions, but not/v1/responses, even though OpenRouter's own Responses endpoint would accept them. Other chat models serve chat completions and the Responses API, and embedding models serve embeddings. - Reported spend matches your OpenRouter invoice. OpenRouter prices each request itself and reports the figure inline, and the gateway bills from that rather than from a stored rate card. That matters because OpenRouter's routing shortcuts can send the same model to different hosts at different prices from one request to the next.
For prerequisites, provider steps, code examples, the routing shortcuts, and troubleshooting, see Using OpenRouter.
Jev
Jev is TypeSafe's System One model. It doesn't generate text: you send it a piece of text and a set of typed questions, and it returns a calibrated answer for each one—a probability, a choice from your options, or a score on your scale.
Three things are specific to it:
- It has its own endpoint. Requests go to
/v1/systemone. Jev models are refused on chat completions, Responses, Messages, and embeddings, and they don't appear in the defaultGET /v1/modelslist. - Model references are
jev/<model>. A provider namedjevservesjev/jev-latest,jev/jev-1.13.0, andjev/jev-preview. The alias names move when TypeSafe ships a release, so pin a versioned name if you tune thresholds against one. - Only input tokens cost money. TypeSafe charges for input tokens and not for output, and the gateway prices requests from the model catalog.
Guardrail policies such as DLP deny System One requests by default, because they can't inspect the body. For prerequisites, provider steps, code examples, and the policy table, see Using Jev.
Zuplo Demo
Zuplo Demo is a free provider that Zuplo operates so you can try the AI Gateway without an account with an LLM provider. Add it like any other provider—see Managing Providers—but skip the API key: the dialog doesn't ask for one, because your gateway authenticates to the demo service itself.
The provider serves three demo persona models as chat completions, at no cost:
zuplodemo/piratezuplodemo/wizardzuplodemo/space-cowboy
Not for production use
Zuplo Demo has daily limits for each Zuplo account—plenty for exploring the gateway and demos, but not for real traffic. Swap in a keyed provider such as OpenAI or Anthropic before going live.
If you need support for additional providers or capabilities, please contact us at support@zuplo.com. We're continually working to add support for more providers based on customer demand.