Model Load Balancing Policy
AI Gateway Policy
This policy is for use with the AI Gateway. See the AI Gateway documentation to learn how to configure and govern AI models with Zuplo.
Spreads AI requests across a list of models. The policy cycles through the list, or keeps requests that share a key, such as a user or a conversation, on the same model. Most requests also get a retry model from the list. A request moves there when the chosen model stays rate limited. Chat completions and embeddings requests also move there when the model returns a server error, can't be reached, or times out.
Configuration
The configuration shows how to configure the policy in the 'policies.json' document.
config/policies.json
Policy Configuration
name<string>- The name of your policy instance. This is used as a reference in your routes.policyType<string>- The identifier of the policy. This is used by the Zuplo UI. Value should beai-gateway-load-balancing-inbound.handler.export<string>- The name of the exported type. Value should beAIGatewayLoadBalancingInboundPolicy.handler.module<string>- The module containing the policy. Value should be$import(@zuplo/runtime).handler.options<object>- The options for this policy. See Policy Options below.
Policy Options
The options for this policy are specified below. All properties are optional unless specifically marked as required.
strategy<string>- How the policy picks a model for each request.round-robincycles through the list in order.stickysends every request with the same key to the same model. Allowed values areround-robin,sticky. Defaults to"round-robin".stickyBy<string>- The key thatstickyuses.useruses the authenticated caller,request.user.sub.conversationuses the conversation's system prompt and its first user message with text.expressionuses the value thatexpressionselects. Required whenstrategyissticky, and not allowed withround-robin. Allowed values areuser,conversation,expression.expression<string>- The value to key on, written like a Metering budget expression, such asrequest.headers.get("x-end-user-id"). Required whenstickyByisexpression, and not allowed otherwise.models(required)<object>- The models to spread requests across, by request type. Set at least one list.completions<string[]>- Models for chat completions, Responses, Messages, and System One requests, asproviderName/model. From 1 through 32 models, each listed once. Provider names aren't case-sensitive, soopenai/gpt-5andOpenAI/gpt-5are the same model. Write each model ID exactly as your provider spells it.embeddings<string[]>- Models for embeddings requests, asproviderName/model. From 1 through 32 models, each listed once. Provider names aren't case-sensitive. Write each model ID exactly as your provider spells it. Every entry must be the same embedding model, such as one model on several accounts.
fallbackTimeoutSeconds<integer>- How long to wait for a model to start responding before retrying the request on another model. Applies to chat completions and embeddings. The timeout ends when the response starts, so it doesn't cut off a long streaming response. A response that doesn't stream starts only when it's complete, so the timeout covers the whole response. From 1 through 300. Defaults to60.
Using the Policy
Use this policy to spread requests across several deployments of a model to combine their rate limits, across providers that serve the same model, or across different models. List the models, and every request the policy handles goes to one of them.
How it works
For each request, the policy:
- Takes the list for the request's type:
completionsfor chat completions, Responses, Messages, and System One, orembeddingsfor embeddings. - Drops models that can't serve the request's endpoint, such as an OpenAI model
on
/v1/messages. The policy decides this by the provider and by whether the model is Claude, not by which APIs each model serves. See Endpoints. - Picks a model with the
strategy, plus a retry model: the model the strategy would pick next. - Hands both to the AI Gateway. The gateway calls the chosen model, and moves to the retry model when that call fails in a way the endpoint retries. See Failures and retries.
The gateway sends the policy's choice to the provider in place of the request's
model. Clients still send a model, because most SDKs require one, but any
value works. The request itself keeps the client's model, so policies that run
later, such as Semantic Cache, still see that value. Because the policy picks
the model for every request it handles, the list also limits which models
clients can reach.
If an earlier policy already chose a model for the request, this policy leaves
that choice alone. Requests of a type with no list, such as embeddings when you
set only completions, pass through unchanged.
Choose a strategy
| Strategy | How it picks a model | Use it to |
|---|---|---|
round-robin | Cycles through the list in order. | Spread load evenly across deployments, accounts, or providers. |
sticky | Sends every request with the same key to the same model. The key is the caller, the conversation, or a value you select. | Keep each caller or end user on one model, or keep each conversation on the account whose prompt cache already holds it. |
Here is how four requests could go with the example configuration, on one gateway instance:
| Request | Caller | round-robin | sticky by user |
|---|---|---|---|
| 1 | app-a | azure-eastus/gpt-5 | azure-eastus/gpt-5 |
| 2 | app-b | azure-westus/gpt-5 | openai/gpt-5 |
| 3 | app-a | azure-westus/gpt-5 | azure-eastus/gpt-5 |
| 4 | app-a | openai/gpt-5 | azure-eastus/gpt-5 |
Round robin moves each app to the next model on each of its requests. Sticky
keeps app-a on azure-eastus/gpt-5, whichever gateway instance serves it.
Examples
Serve one model from several providers
Claude Sonnet is available directly from Anthropic and through Amazon Bedrock. Cycle between them to use both quotas. Use a Bedrock Mantle provider for the Bedrock entry, because the policy skips models on Bedrock Runtime providers:
Code
Keep conversations on one deployment
Providers cache the prompt prefixes they've already processed, and cached input tokens cost less and return faster. Each account or deployment has its own cache, so a conversation that moves between deployments loses the benefit:
Code
This works for clients that send the whole conversation on every turn, as Chat
Completions clients do. A Responses API client that continues a stored response
with previous_response_id sends only the new turn, so with this list, its
requests to read or continue a response fail. Use sticky routing by user or
expression for those clients. See Stored responses.
Keep each customer on one model
When each of your customers calls through its own app, sticky routing by user keeps every customer on one model while customers spread across the list:
Code
Balance embeddings
Code
Use one embedding model per list
Every entry in the embeddings list must be the same embedding model. Vectors
from different models can't be compared, so mixing models breaks similarity
search.
Round robin
Round robin gives every model in the list an equal share of requests. Each gateway instance keeps a separate place in the list for each app, starting from a random point. Instances never have to coordinate, and a busy app can't skew a quiet app's rotation. Across all instances, requests spread evenly, but two requests in a row can still go to the same model.
Sticky routing
Sticky routing uses a hash, not a stored mapping, so it gives the same answer on every gateway instance. For each request, the policy scores every model by hashing the sticky key together with the model name, and picks the model with the highest score. The same key always produces the same scores, so it always gets the same model.
| Model | Score for app-a | Result |
|---|---|---|
azure-eastus/gpt-5 | 0.74 | Chosen |
azure-westus/gpt-5 | 0.72 | Retry model |
openai/gpt-5 | 0.25 |
- The retry model is the one with the second-highest score, so it's stable too.
- Adding a model moves only the keys it now wins. Removing a model moves only the keys it was winning. Every other key stays where it was.
Requests without a key go round robin:
stickyBy | A request has no key when |
|---|---|
user | It has no authenticated user |
conversation | It has no user message with text, such as an embeddings or System One request |
expression | The expression resolves to nothing, to an empty string, or to an unusable value |
Stored-response requests don't go round robin when the models that store
responses have more than one provider name. They fail with 400 unless a user
or expression key picks their model, so under conversation they fail even
with a key. See Stored responses.
Sticky by user
"stickyBy": "user" keys on request.user.sub, the identity your
authentication policy sets. It works like an expression of request.user.sub.
With AI Gateway API keys, the user is the app
AI Gateway Authentication sets request.user.sub to the app's name. Sticky
routing by user then keeps each app on one model, and spreads load only when
many apps share the policy. To keep each end user on one model, have the app
send an end-user ID and select it with an expression.
On User App routes, where callers use personal API keys, request.user.sub is
the key owner's user ID, so each person stays on one model.
Sticky by conversation
"stickyBy": "conversation" keys on the conversation's system prompt and its
first user message with text. A client that sends the whole conversation on
every turn, as Chat Completions and Messages clients do, sends the same key on
every turn. On the Responses API, the system prompt is instructions plus any
system or developer items before the first user message. Keeping a conversation
on one model keeps it on the account whose prompt cache already holds its
earlier turns.
Only text counts, including plain-text documents. Conversations that open with the same system prompt and the same first message share a model, and images and other files don't make them different. A conversation that opens with an image on its own goes round robin until a turn adds text. If your clients open every conversation the same way, send a conversation ID in a header and key on it with an expression instead.
A Responses API request that continues a stored response with
previous_response_id or conversation sends only the new turn, so its key
changes on every turn. See Stored responses.
The policy hashes this text in memory. It never stores or logs it.
Sticky by expression
"stickyBy": "expression" keys on one value that expression selects from the
request or its context. Expressions use the same syntax as
Metering budget expressions,
so a value you budget on can also keep requests on one model. For example, to
keep each end user on one model when your app sends an end-user ID in a header:
Code
In JSON, escape the double quotes inside an expression, as the example does. Common expressions:
| Expression | Selects |
|---|---|
request.headers.get("<name>") | One request header, such as an end-user ID your app sends |
request.user.data.<property> | A property of the caller's metadata |
request.params.<name> | A path parameter from the matched route |
context.custom.<property> | A value your own policy puts on context.custom |
See Supported expressions in the Metering docs for the full list. An expression is a selector, not code, so it can't compare values, call string methods, or read the request body.
- The value follows Metering's value rules. It must be a string or a safe
integer, and
-0doesn't count. The gateway converts a string to well-formed NFC Unicode, then rejects it when it's longer than 256 UTF-8 bytes or contains a control character, U+2028, or U+2029. - Values are case-sensitive, so
Acmeandacmecan land on different models. Two strings that normalize to the same NFC text share a model. - When the expression resolves to nothing, to an empty string, or to a value
these rules reject, such as an object, a fractional number, or an overlong
string, the request goes round robin. A stored-response request fails with
400instead when the models that store responses have more than one provider name. The request log says why. - Per-request identifiers, such as
context.requestId, aren't allowed, because a key that never repeats can't keep requests together. - An expression outside the syntax is an options error, and the error shows where parsing stopped.
The policy hashes the value in memory. It never stores or logs it.
Failures and retries
Each request gets a retry model, except stored-response requests,
/v1/messages/count_tokens requests, /v1/systemone requests, and requests
that only one model in the list can serve. The AI Gateway treats the retry model
like any other backup model. When it moves there depends on the endpoint:
| When the chosen model | Chat completions and embeddings | Responses and Messages |
|---|---|---|
Returns 429 | Retries it twice, then moves to the retry model | Retries it twice, then moves to the retry model |
Returns a 5xx, 408, or 425 status, or can't be reached | Moves to the retry model | Fails the request |
Doesn't start responding within fallbackTimeoutSeconds | Moves to the retry model | Keeps waiting |
Returns any other error, such as 400 or 401 | Fails the request | Fails the request |
Responses and Messages requests use the provider's native API. On those endpoints the AI Gateway moves to another model only for rate limits, as it does without this policy.
- Each request tries at most two models. The gateway retries a rate-limited
model twice, after 250 ms and then 500 ms. If every attempt returns
429, the request makes six calls, three to each model, before the gateway returns429. - Once a model starts streaming a response, the request stays on that model.
- A request that doesn't stream gets its response only when the model finishes,
so on chat completions and embeddings
fallbackTimeoutSecondslimits the whole response. A model that's still working when it runs out moves the request to the retry model, which starts over with twice the time, and the request fails if that runs out too. RaisefallbackTimeoutSecondsfor slow models, such as reasoning models, or stream the response. When Fallback Model runs after this policy with afallback, raise itsfallbackTimeoutSecondstoo. Its timeout replaces this policy's on requests that get itsfallback. See With Fallback Model. - When only one model in the list can serve a request, this policy sets no retry
model and no timeout. On
/v1/responsesand/v1/messages, a list that mixes providers or vendors often has only one, so list at least two models for each API your clients use. Fallback Model can also add a backup. See With Fallback Model. - The policy doesn't remember failures between requests. While a model is down,
every request that picks it tries it first: chat completions and embeddings
requests then move to the retry model, and Responses and Messages requests
fail. With
stickyrouting, the same callers pick the failing model every time.
Endpoints
| Endpoint | Behavior |
|---|---|
/v1/chat/completions | Models the AI Gateway sends Chat Completions requests to are eligible, and the gateway translates the request for each provider. System One models are skipped. Claude on Vertex AI isn't skipped, and requests that pick it fail with 400. |
/v1/responses | Only models the AI Gateway sends Responses API requests to are eligible: OpenAI models, Azure AI and OpenRouter models other than Claude, and Bedrock Mantle models other than Claude. Other providers are skipped, even when their own API serves the Responses API. |
/v1/messages, /v1/messages/count_tokens | Only models the AI Gateway sends Anthropic Messages API requests to are eligible: Anthropic models, and Claude on Azure AI, Bedrock Mantle, Vertex AI, and OpenRouter. count_tokens also skips providers that don't count tokens, such as OpenRouter. |
/v1/systemone | Only TypeSafe System One (Jev) models are eligible, so you can balance several Jev models or TypeSafe accounts. Jev through OpenRouter isn't eligible. These requests get no retry model. |
/v1/embeddings | Uses the embeddings list. |
GET or DELETE /v1/responses/{id}, and GET /v1/responses/{id}/input_items | Routed to the account that holds the response, when the policy can tell which one it is. See Stored responses. |
GET /v1/models | Passes through unchanged. |
Bedrock Runtime (/model/{modelId}/…) | Not supported. An app whose policy chain includes this policy rejects Bedrock Runtime requests. A model on a Bedrock Runtime provider can't serve any /v1 endpoint, so the policy skips it. |
If no model in the list can serve the request's endpoint, the policy returns
400. The message names the endpoint, and for an endpoint that uses a
provider's own API, such as /v1/messages, the API it requires.
These rules depend on the provider, and on whether the model is Claude. The
policy doesn't check which APIs each model serves. A request that picks a model
on an API the model doesn't serve fails, such as a /v1/chat/completions
request that picks a model that serves only the Responses API.
On /v1/responses:
- Bedrock Mantle serves the Responses API for only some of its models other than
Claude. Requests that pick a Mantle model without Responses support fail. To
check a model, read the APIs supported for the
bedrock-mantleendpoint on its AWS model card. The AI Gateway calls that endpoint, and thebedrock-runtimeendpoint's list can differ. - OpenRouter stores no responses. When the list also has a model that does,
OpenRouter gets only creates that set
storetofalse. See Stored responses.
Stored responses
A stored response lives on the provider account that created it. Requests that
read or continue one must reach that account: retrieving a response, listing its
input items, deleting it, and creating a response with previous_response_id, a
conversation, or an item_reference input item.
| When | Stored-response requests |
|---|---|
| Every model in the list that stores responses has the same provider name | Go to that provider |
strategy is sticky by user or expression, and the request has a key | Go to the model the key picks, where that caller's responses were created |
| Anything else | Fail with 400, and the request log names the fix |
- The provider name is the part before the slash, compared without regard to
case.
openai/gpt-5andopenai/gpt-5-minishare one, butazure-eastus/gpt-5andazure-westus/gpt-5are two. - Stored-response requests get no retry model, because no other account holds
the response. Fallback Model placed after this policy still adds its
fallback, so a continuation that stays rate limited moves to the fallback, and fails there if the fallback is on another account. - A provider that stores nothing, such as OpenRouter, never gets stored-response
requests. It gets a create only when the create sets
storetofalse, or when no provider in the list stores responses, because a create stores its response unless it opts out. - A create that moves to the retry model stores its response on the retry model's account. If that's another account, later requests for the response go to the chosen model's account and fail.
- The retry model, a
fallback, or aquotaFallbackis on another account when its provider name differs from the chosen model's. It's also on another account when AI Gateway Authentication passes callers' own API keys through, because the chosen model uses the caller's key and the others use your configured key. - The
400tells the caller only that this API can't read or continue the stored response. The fix is yours to make, so the request log carries it. conversationkeys can't route them to one of several provider names. A retrieve sends no messages, and a request withprevious_response_idsends only the new turn. With one provider name, each continuation is keyed on its new turn, so it can run on a different model than the turn before.- Changing the list can move a caller to another model. Responses stored before the change stay on the old account.
Policy order
Place Model Load Balancing after authentication and before the policies that read or change model routing:
Code
This chain shows the order only. Entries without options use the options
declared in config/policies.json. See
Configure lists per app.
| Policy | With Model Load Balancing |
|---|---|
| Authentication | Place it before, so request.user.sub is set for sticky routing by user. |
| Model Override | Not needed, because this policy picks the model for every request it handles. If Model Override runs first, its model wins: always with force, and with default when a request omits model, as stored-response reads do. Placed after, it does nothing for request types this policy has a list for. |
| Model Filtering | Not needed for request types this policy has a list for, because the list already limits the models clients can reach. If Model Filtering runs first, it routes or refuses each request by the model the client asked for, and this policy does nothing. Placed after, it doesn't check this policy's choices. |
| Fallback Model | Place it after, where it changes the retry model, the timeout, and where over-quota requests go. Placed before, it does nothing. See With Fallback Model. |
| Smart Router | When it routes a request, its model wins, whichever runs first. The request loses this policy's retry model, and Smart Router doesn't apply this policy's endpoint or stored-response rules. Fallback Model placed after Smart Router still adds its fallback and quotaFallback. Placed before Smart Router, they're lost too. |
| Metering | Place it after this policy and Fallback Model, so a request over its quota can switch to Fallback Model's quotaFallback. Costs and usage use the model that served the request. |
| Semantic Cache | Place it after. Its cache key uses the request's own model, not this policy's choice, so with a list of different models a cached answer can come from any of them. Placed before, it serves a cache hit without this policy when the request's model is a valid model, and skips the hit otherwise. |
With Fallback Model
Fallback Model placed after this policy changes where failed and over-quota requests go:
| Requests | Fallback Model's fallback | Its quotaFallback, over quota |
|---|---|---|
| Chat completions and embeddings | Becomes the retry model, and its fallbackTimeoutSeconds, 60 when unset, replaces this policy's. | Gets the request. A quotaFallback that the AI Gateway refuses on chat completions, such as Claude on Vertex AI, fails the request with 400. |
| Responses and Messages creates | Becomes the retry model if it can serve the endpoint. Otherwise the request has no retry model. | Gets the request if it can serve the endpoint. Otherwise the gateway returns 429. |
/v1/messages/count_tokens | Never tried. | Gets the request if it can serve the endpoint. Otherwise the gateway returns 429. |
/v1/systemone, and requests that retrieve, delete, or list the input items of a stored response | Never tried. | Never used. The gateway returns 429. |
- Requests this policy sends to the
fallbackmodel itself keep this policy's retry model and timeout. Requests it sends to thequotaFallbackmodel itself get no quota fallback, so when they're over quota, the gateway returns429. Use aquotaFallbackthat isn't in the list. - When a
quotaFallbackcall fails in a way the endpoint retries, the request moves to the retry model, so an over-quota request can still run on a model in the list. To keep over-quota requests off the list, give Fallback Model afallbackthat isn't in the list. To keep failed requests inside the list, give Fallback Model only aquotaFallback. - No model serves both
/v1/responsesand/v1/messages, so when your clients call both, afallbackor aquotaFallbackcovers at most one of them. - Over quota, a create that continues a stored response goes to the
quotaFallback, and fails there if thequotaFallbackis on another account. Other creates over quota go there too, so a response they store lives on thequotaFallback's account. If that's another account, requests for the response fail once they're no longer over quota. See Stored responses. - For embeddings, set
fallbackto the same embedding model as the list. With a one-model list, use that model on another provider name, because afallbackequal to the chosen model is skipped. - Unless you've changed them, Fallback Model's declared options set
openai/gpt-4o-minias thecompletionsfallback. A Fallback Model entry withoutoptions, like the one in the chain above, makes that model the retry model. On a gateway that can't useopenai/gpt-4o-mini, such as one with no provider namedopenai, every request that Fallback Model adds it to fails with500, even requests that never try a backup. Give the entry its ownoptions.
Configure lists per app
Declare the policy in config/policies.json with a default list. An app can
supply its own options in its chain entry. Entry options replace the declared
options entirely:
Code
Unless you've changed them, the declared options are a one-model list,
openai/gpt-4o-mini. A chain entry without options then sends every chat
completions and Responses request to that model, and fails Messages and System
One requests with 400, because that model can't serve them.
The gateway validates chain options when the policy runs. Invalid options fail requests with an error that names the field to fix. So does a model the gateway can't use, such as an unknown provider or an inactive model, on every request that picks it as the chosen or the retry model. A provider whose API key isn't on the gateway yet fails those requests with an error that says the provider isn't ready, and the request log names the provider and its environment variable.
Monitoring
AI Gateway analytics attribute each request to the model that served it. Logging plugins also receive these properties on the request's log entries after the policy picks a model:
| Property | Value |
|---|---|
aiLoadBalancingModel | The model the policy chose |
aiLoadBalancingRetryModel | The retry model, when there is one |
The properties record this policy's picks. They don't change when Fallback Model replaces the retry model or Smart Router replaces the choice.
Limitations
- Each list can have up to 32 models, each listed once. The policy compares
entries without regard to case, so
openai/gpt-5andOpenAI/gpt-5count as one model. Provider names work in any case, but write each model ID exactly as your provider spells it. A model ID in another case, such asopenai/GPT-5, can fail to route. - Each request tries at most two models.
- An expression can be up to 1024 bytes and 32 property segments.
Lists that mix providers or models have limits of their own:
| When | What happens |
|---|---|
A list mixes providers or vendors, on /v1/responses or /v1/messages | Only models the AI Gateway sends that API's requests to are eligible, and that depends on the model, not only its provider. On Azure AI, for example, Claude models are eligible on /v1/messages, and GPT models on /v1/responses. See Endpoints. A mixed list often leaves one model and no retry model, so list at least two models for each API your clients use. |
A list mixes vendors, on /v1/chat/completions | Each pick goes through its provider's translation, which drops what that provider doesn't support. Claude models ignore response_format and seed, use a max_tokens of 1,024 unless the request sets max_tokens, read developer messages as user messages, and fail requests that carry images, audio, or files. For the same behavior on every pick, list deployments of one model. |
| The models accept different parameters or context lengths | A request that one model rejects fails whenever the policy picks that model. The retry model isn't tried. |
AI Gateway Authentication passes callers' own API keys through (credentialPassthrough) | The caller's key goes to whichever model the policy picks, so list only models that accept it. Every model that can be the retry model also needs its own configured key, and requests that move there run on your provider account. |
| You want different lists for Chat Completions, Responses, and Messages | One completions list serves them all. A second Model Load Balancing policy in the chain doesn't run for request types the first one has a list for. |
Read more about how policies work