Providers

AI Providers

Zuplo's AI Gateway supports integration with various AI providers, allowing you to leverage different models and services for your apps.

Supported Providers

Zuplo currently supports the following AI providers:

  • OpenAI
  • Anthropic
  • Google
  • Mistral
  • xAI (Grok)
  • Azure AI—Azure OpenAI and Microsoft Foundry resources, serving Claude and every Foundry model family from your own Azure subscription
  • Bedrock Mantle—Amazon Bedrock's compatible-APIs endpoint, serving Claude models and models from many other vendors
  • Bedrock Runtime—Amazon Bedrock's native API, for code that already calls Bedrock with an AWS SDK
  • Vertex AI—Google Cloud's managed model platform, serving Gemini and Model Garden models on your own Google Cloud project
  • OpenRouter—hundreds of models from every major vendor behind a single API key
  • Jev—TypeSafe's System One model, which answers typed questions about text instead of generating it
  • Zuplo Demo—a free, keyless provider for trying the gateway
  • OpenAI-compatible Custom Providers (such as Qwen, Kimi, etc)

The following capabilities are supported across providers:

ProviderChat CompletionsEmbeddingsResponsesMessages
OpenAI✅✅✅❌
Anthropic✅❌❌✅
Google✅✅❌❌
Mistral✅✅❌❌
xAI✅✅❌❌
Azure AI✅✅✅✅
Bedrock Mantle✅❌✅✅
Bedrock Runtime❌❌❌❌
Vertex AI✅✅❌✅
OpenRouter✅✅✅✅
Jev❌❌❌❌
Zuplo Demo✅❌❌❌
OpenAI-compatible (Custom)✅✅❌❌

Responses is the OpenAI Responses API (/v1/responses), and Messages is the native Anthropic Messages API (/v1/messages). See the Universal API for the endpoint each capability maps to.

Bedrock Mantle's capabilities depend on the model family: Claude models serve Messages (plus chat completions through translation), while its other models serve chat completions and—per model—Responses. See Using Bedrock Mantle.

Bedrock Runtime and Jev serve none of these capabilities, and the table above says so. Bedrock Runtime serves Amazon Bedrock's native API instead, at its own paths, so your existing AWS SDK code keeps working with the gateway in front of it. See Using Bedrock Runtime for the operations it serves, and add Bedrock Mantle alongside it if you also want Bedrock models on the Universal API.

Jev serves TypeSafe's System One API instead, at /v1/systemone. You send it text and typed questions, and it returns typed answers rather than generated text. See Using Jev.

Azure AI's capabilities split by model family, and within the OpenAI-compatible family they split per deployed model: a chat model serves chat completions and— per model and version—Responses, while an embedding model serves embeddings and neither of the others. Claude serves Messages (plus chat completions through translation) and needs a Microsoft Foundry resource, since Claude can't be deployed on an Azure OpenAI resource. Azure also addresses models by deployment name rather than by published model ID, so a deployment named differently from the model it serves needs mapping for the gateway to price it. See Using Azure AI.

A custom provider must serve chat completions and embeddings under a /v1 path segment on its API URL, and you enter that URL as an origin root—without the /v1 suffix vendors usually publish. See Custom Providers for the exact contract.

Apps reference a provider's models as providerName/model—for example openai/gpt-6-luna or anthropic/claude-sonnet-5—where providerName is the name you give the provider configuration. See the Universal API.

Azure AI

Azure AI serves models from your own Azure resource. One provider configuration covers both kinds of resource, and which capabilities apply depends on the resource and the model:

  • An Azure OpenAI resource deploys OpenAI models and serves chat completions, embeddings, and the Responses API.
  • A Microsoft Foundry resource serves those plus every Foundry model family—Grok, DeepSeek, Llama, Mistral, Phi, Kimi and more—and serves Claude models on the native Anthropic Messages API, with chat completions through the gateway's translation.

Two things make Azure different from every other provider:

  • It addresses deployments, not published model IDs. You name deployments when you create them in Azure, and that name is what goes in the model field. A deployment whose name differs from the model it serves needs a mapping, or the gateway can't price its usage.
  • Its endpoint host embeds your resource name. The provider dialog asks for an Azure Resource Name and a host family rather than a full URL, and it takes a resource key—not a Microsoft Entra ID token.

For prerequisites, setup steps, code examples, and troubleshooting, see Using Azure AI.

Bedrock Mantle

Bedrock Mantle is Amazon Bedrock's compatible-APIs endpoint. One regional endpoint and one long-term Bedrock API key serve two model families, and which capabilities apply depends on the model you call:

  • Claude models (for example anthropic.claude-sonnet-5) serve the native Anthropic Messages API, and chat completions through the gateway's translation.
  • OpenAI-compatible models—everything else Mantle serves, from open-weight models such as openai.gpt-oss-120b to frontier models such as openai.gpt-5.6-sol—serve chat completions and, per model, the OpenAI Responses API.

Mantle endpoints are regional, so the provider dialog asks for an AWS Region instead of an endpoint URL, and it accepts only long-term Bedrock API keys, which start with ABSK. For prerequisites, setup steps, code examples, and troubleshooting, see Using Bedrock Mantle.

Bedrock Runtime

Bedrock Runtime is Amazon Bedrock's native API—the one the AWS SDKs call. Unlike every other provider, it doesn't serve the Universal API: it serves Bedrock's own operations (Converse, ConverseStream, InvokeModel, and InvokeModelWithResponseStream) at Bedrock's own paths.

That makes it the provider to choose when you already have code calling Bedrock with an AWS SDK. You point the client's endpoint at your app and give it the app's API key; every call site stays as it is, and the gateway adds authentication, model filtering, and usage metering in front of your own AWS account.

It authenticates to AWS with a long-term IAM access key pair rather than a Bedrock API key, so it needs its own provider configuration even if you also use Bedrock Mantle on the same account. For prerequisites, setup steps, code examples, and troubleshooting, see Using Bedrock Runtime.

Vertex AI

Vertex AI is Google Cloud's managed model platform. One provider configuration serves both the Gemini family and Vertex's Model Garden partner models—DeepSeek, Qwen, GLM, Kimi, MiniMax, Gemma, and GPT-OSS—from your own Google Cloud project, on your Google Cloud billing.

Two things make its setup different from every other provider:

  • It authenticates with a service account, not an API key. Vertex's prediction endpoints reject API keys, so the dialog asks for a service account JSON key file and the gateway exchanges it for short-lived access tokens.
  • It needs a Google Cloud project as well as a location. Vertex endpoints are per-location, and the project isn't part of the endpoint, so the dialog asks for both.

Every Vertex model ID is publisher-qualified, so a model reference has two slashes: a provider named vertexai serves vertexai/google/gemini-3.7-flash, vertexai/qwen/qwen3-coder-480b-a35b-instruct-maas, and vertexai/anthropic/claude-haiku-4-5@20251001. Model availability varies by location—most Model Garden models are served only from the global location, and Google's regional endpoints serve Claude Sonnet 4.6 and earlier only, so global is the location that serves every Claude model. Vertex's Claude models serve the native Anthropic Messages API (/v1/messages) and nothing else: they aren't available on /v1/chat/completions, and each one must be enabled in Model Garden for your Google Cloud project before it answers.

For prerequisites, Google Cloud setup, provider steps, code examples, and troubleshooting, see Using Vertex AI.

OpenRouter

OpenRouter is a marketplace, not a single vendor: one API key reaches hundreds of models from OpenAI, Anthropic, Google, Meta, DeepSeek, Qwen, xAI and many more, and OpenRouter picks an upstream host for each request.

Three things are specific to it:

  • Model references have two slashes. Every OpenRouter model ID is vendor-qualified, so a provider named openrouter serves openrouter/anthropic/claude-sonnet-5 and openrouter/openai/gpt-6-sol. This is the same shape Vertex AI uses.
  • Claude models are on the Messages API. Models whose ID begins anthropic/ serve the native Anthropic Messages API (/v1/messages) and a translated /v1/chat/completions, but not /v1/responses, even though OpenRouter's own Responses endpoint would accept them. Other chat models serve chat completions and the Responses API, and embedding models serve embeddings.
  • Reported spend matches your OpenRouter invoice. OpenRouter prices each request itself and reports the figure inline, and the gateway bills from that rather than from a stored rate card. That matters because OpenRouter's routing shortcuts can send the same model to different hosts at different prices from one request to the next.

For prerequisites, provider steps, code examples, the routing shortcuts, and troubleshooting, see Using OpenRouter.

Jev

Jev is TypeSafe's System One model. It doesn't generate text: you send it a piece of text and a set of typed questions, and it returns a calibrated answer for each one—a probability, a choice from your options, or a score on your scale.

Three things are specific to it:

  • It has its own endpoint. Requests go to /v1/systemone. Jev models are refused on chat completions, Responses, Messages, and embeddings, and they don't appear in the default GET /v1/models list.
  • Model references are jev/<model>. A provider named jev serves jev/jev-latest, jev/jev-1.13.0, and jev/jev-preview. The alias names move when TypeSafe ships a release, so pin a versioned name if you tune thresholds against one.
  • Only input tokens cost money. TypeSafe charges for input tokens and not for output, and the gateway prices requests from the model catalog.

Guardrail policies such as DLP deny System One requests by default, because they can't inspect the body. For prerequisites, provider steps, code examples, and the policy table, see Using Jev.

Zuplo Demo

Zuplo Demo is a free provider that Zuplo operates so you can try the AI Gateway without an account with an LLM provider. Add it like any other provider—see Managing Providers—but skip the API key: the dialog doesn't ask for one, because your gateway authenticates to the demo service itself.

The provider serves three demo persona models as chat completions, at no cost:

  • zuplodemo/pirate
  • zuplodemo/wizard
  • zuplodemo/space-cowboy

Not for production use

Zuplo Demo has daily limits for each Zuplo account—plenty for exploring the gateway and demos, but not for real traffic. Swap in a keyed provider such as OpenAI or Anthropic before going live.

If you need support for additional providers or capabilities, please contact us at support@zuplo.com. We're continually working to add support for more providers based on customer demand.

Last modified on