
# Model Load Balancing Policy

:::note{title="AI Gateway Policy"}

This policy is for use with the [AI Gateway](/docs/ai-gateway/overview). See
the AI Gateway documentation to learn how to configure and govern AI models
with Zuplo.

:::

Spreads AI requests across a list of models. The policy cycles through the list,
or keeps requests that share a key, such as a user or a conversation, on the
same model. Most requests also get a retry model from the list. A request moves
there when the chosen model stays rate limited. Chat completions and embeddings
requests also move there when the model returns a server error, can't be
reached, or times out.

## Configuration

The configuration shows how to configure the policy in the 'policies.json' document.

```json title="config/policies.json"
{
  "name": "my-ai-gateway-load-balancing-inbound-policy",
  "policyType": "ai-gateway-load-balancing-inbound",
  "handler": {
    "export": "AIGatewayLoadBalancingInboundPolicy",
    "module": "$import(@zuplo/runtime)",
    "options": {
      "strategy": "round-robin",
      "models": {
        "completions": [
          "azure-eastus/gpt-5",
          "azure-westus/gpt-5",
          "openai/gpt-5"
        ]
      }
    }
  }
}
```

### Policy Configuration

- `name` <code className="text-green-600">&lt;string&gt;</code> - The name of your policy instance. This is used as a reference in your routes.
- `policyType` <code className="text-green-600">&lt;string&gt;</code> - The identifier of the policy. This is used by the Zuplo UI. Value should be `ai-gateway-load-balancing-inbound`.
- `handler.export` <code className="text-green-600">&lt;string&gt;</code> - The name of the exported type. Value should be `AIGatewayLoadBalancingInboundPolicy`.
- `handler.module` <code className="text-green-600">&lt;string&gt;</code> - The module containing the policy. Value should be `$import(@zuplo/runtime)`.
- `handler.options` <code className="text-green-600">&lt;object&gt;</code> - The options for this policy. [See Policy Options](#policy-options) below.

### Policy Options

The options for this policy are specified below. All properties are optional unless specifically marked as required.

- `strategy` <code className="text-green-600">&lt;string&gt;</code> - How the policy picks a model for each request. `round-robin` cycles through the list in order. `sticky` sends every request with the same key to the same model. Allowed values are `round-robin`, `sticky`. Defaults to `"round-robin"`.
- `stickyBy` <code className="text-green-600">&lt;string&gt;</code> - The key that `sticky` uses. `user` uses the authenticated caller, `request.user.sub`. `conversation` uses the conversation's system prompt and its first user message with text. `expression` uses the value that `expression` selects. Required when `strategy` is `sticky`, and not allowed with `round-robin`. Allowed values are `user`, `conversation`, `expression`.
- `expression` <code className="text-green-600">&lt;string&gt;</code> - The value to key on, written like a [Metering budget expression](https://zuplo.com/docs/policies/ai-gateway-metering-inbound#budget-expressions), such as `request.headers.get("x-end-user-id")`. Required when `stickyBy` is `expression`, and not allowed otherwise.
- `models` **(required)** <code className="text-green-600">&lt;object&gt;</code> - The models to spread requests across, by request type. Set at least one list.
  - `completions` <code className="text-green-600">&lt;string[]&gt;</code> - Models for chat completions, Responses, Messages, and System One requests, as `providerName/model`. From 1 through 32 models, each listed once. Provider names aren't case-sensitive, so `openai/gpt-5` and `OpenAI/gpt-5` are the same model. Write each model ID exactly as your provider spells it.
  - `embeddings` <code className="text-green-600">&lt;string[]&gt;</code> - Models for embeddings requests, as `providerName/model`. From 1 through 32 models, each listed once. Provider names aren't case-sensitive. Write each model ID exactly as your provider spells it. Every entry must be the same embedding model, such as one model on several accounts.
- `fallbackTimeoutSeconds` <code className="text-green-600">&lt;integer&gt;</code> - How long to wait for a model to start responding before retrying the request on another model. Applies to chat completions and embeddings. The timeout ends when the response starts, so it doesn't cut off a long streaming response. A response that doesn't stream starts only when it's complete, so the timeout covers the whole response. From 1 through 300. Defaults to `60`.

## Using the Policy

Use this policy to spread requests across several deployments of a model to
combine their rate limits, across providers that serve the same model, or across
different models. List the models, and every request the policy handles goes to
one of them.

## How it works

For each request, the policy:

1. Takes the list for the request's type: `completions` for chat completions,
   Responses, Messages, and System One, or `embeddings` for embeddings.
2. Drops models that can't serve the request's endpoint, such as an OpenAI model
   on `/v1/messages`. The policy decides this by the provider and by whether the
   model is Claude, not by which APIs each model serves. See
   [Endpoints](#endpoints).
3. Picks a model with the `strategy`, plus a retry model: the model the strategy
   would pick next.
4. Hands both to the AI Gateway. The gateway calls the chosen model, and moves
   to the retry model when that call fails in a way the endpoint retries. See
   [Failures and retries](#failures-and-retries).

The gateway sends the policy's choice to the provider in place of the request's
`model`. Clients still send a `model`, because most SDKs require one, but any
value works. The request itself keeps the client's `model`, so policies that run
later, such as Semantic Cache, still see that value. Because the policy picks
the model for every request it handles, the list also limits which models
clients can reach.

If an earlier policy already chose a model for the request, this policy leaves
that choice alone. Requests of a type with no list, such as embeddings when you
set only `completions`, pass through unchanged.

## Choose a strategy

| Strategy      | How it picks a model                                                                                                     | Use it to                                                                                                                |
| ------------- | ------------------------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------ |
| `round-robin` | Cycles through the list in order.                                                                                        | Spread load evenly across deployments, accounts, or providers.                                                           |
| `sticky`      | Sends every request with the same key to the same model. The key is the caller, the conversation, or a value you select. | Keep each caller or end user on one model, or keep each conversation on the account whose prompt cache already holds it. |

Here is how four requests could go with the example configuration, on one
gateway instance:

| Request | Caller  | `round-robin`        | `sticky` by `user`   |
| ------- | ------- | -------------------- | -------------------- |
| 1       | `app-a` | `azure-eastus/gpt-5` | `azure-eastus/gpt-5` |
| 2       | `app-b` | `azure-westus/gpt-5` | `openai/gpt-5`       |
| 3       | `app-a` | `azure-westus/gpt-5` | `azure-eastus/gpt-5` |
| 4       | `app-a` | `openai/gpt-5`       | `azure-eastus/gpt-5` |

Round robin moves each app to the next model on each of its requests. Sticky
keeps `app-a` on `azure-eastus/gpt-5`, whichever gateway instance serves it.

## Examples

### Serve one model from several providers

Claude Sonnet is available directly from Anthropic and through Amazon Bedrock.
Cycle between them to use both quotas. Use a Bedrock Mantle provider for the
Bedrock entry, because the policy skips models on Bedrock Runtime providers:

```json
{
  "models": {
    "completions": [
      "anthropic/claude-sonnet-4-6",
      "bedrockmantle/anthropic.claude-sonnet-4-6"
    ]
  }
}
```

### Keep conversations on one deployment

Providers cache the prompt prefixes they've already processed, and cached input
tokens cost less and return faster. Each account or deployment has its own
cache, so a conversation that moves between deployments loses the benefit:

```json
{
  "strategy": "sticky",
  "stickyBy": "conversation",
  "models": {
    "completions": [
      "azure-eastus/gpt-5",
      "azure-westus/gpt-5",
      "azure-swedencentral/gpt-5"
    ]
  }
}
```

This works for clients that send the whole conversation on every turn, as Chat
Completions clients do. A Responses API client that continues a stored response
with `previous_response_id` sends only the new turn, so with this list, its
requests to read or continue a response fail. Use sticky routing by `user` or
`expression` for those clients. See [Stored responses](#stored-responses).

### Keep each customer on one model

When each of your customers calls through its own app, sticky routing by user
keeps every customer on one model while customers spread across the list:

```json
{
  "strategy": "sticky",
  "stickyBy": "user",
  "models": {
    "completions": ["openai/gpt-5", "anthropic/claude-sonnet-4-6"]
  }
}
```

### Balance embeddings

```json
{
  "models": {
    "embeddings": [
      "openai/text-embedding-3-large",
      "azure-eastus/text-embedding-3-large"
    ]
  }
}
```

:::caution{title="Use one embedding model per list"}

Every entry in the `embeddings` list must be the same embedding model. Vectors
from different models can't be compared, so mixing models breaks similarity
search.

:::

## Round robin

Round robin gives every model in the list an equal share of requests. Each
gateway instance keeps a separate place in the list for each app, starting from
a random point. Instances never have to coordinate, and a busy app can't skew a
quiet app's rotation. Across all instances, requests spread evenly, but two
requests in a row can still go to the same model.

## Sticky routing

Sticky routing uses a hash, not a stored mapping, so it gives the same answer on
every gateway instance. For each request, the policy scores every model by
hashing the sticky key together with the model name, and picks the model with
the highest score. The same key always produces the same scores, so it always
gets the same model.

| Model                | Score for `app-a` | Result      |
| -------------------- | ----------------- | ----------- |
| `azure-eastus/gpt-5` | 0.74              | Chosen      |
| `azure-westus/gpt-5` | 0.72              | Retry model |
| `openai/gpt-5`       | 0.25              |             |

- The retry model is the one with the second-highest score, so it's stable too.
- Adding a model moves only the keys it now wins. Removing a model moves only
  the keys it was winning. Every other key stays where it was.

Requests without a key go round robin:

| `stickyBy`     | A request has no key when                                                       |
| -------------- | ------------------------------------------------------------------------------- |
| `user`         | It has no authenticated user                                                    |
| `conversation` | It has no user message with text, such as an embeddings or System One request   |
| `expression`   | The expression resolves to nothing, to an empty string, or to an unusable value |

Stored-response requests don't go round robin when the models that store
responses have more than one provider name. They fail with `400` unless a `user`
or `expression` key picks their model, so under `conversation` they fail even
with a key. See [Stored responses](#stored-responses).

### Sticky by user

`"stickyBy": "user"` keys on `request.user.sub`, the identity your
authentication policy sets. It works like an expression of `request.user.sub`.

:::note{title="With AI Gateway API keys, the user is the app"}

AI Gateway Authentication sets `request.user.sub` to the app's name. Sticky
routing by user then keeps each app on one model, and spreads load only when
many apps share the policy. To keep each end user on one model, have the app
send an end-user ID and select it with an [expression](#sticky-by-expression).
On User App routes, where callers use personal API keys, `request.user.sub` is
the key owner's user ID, so each person stays on one model.

:::

### Sticky by conversation

`"stickyBy": "conversation"` keys on the conversation's system prompt and its
first user message with text. A client that sends the whole conversation on
every turn, as Chat Completions and Messages clients do, sends the same key on
every turn. On the Responses API, the system prompt is `instructions` plus any
system or developer items before the first user message. Keeping a conversation
on one model keeps it on the account whose prompt cache already holds its
earlier turns.

Only text counts, including plain-text documents. Conversations that open with
the same system prompt and the same first message share a model, and images and
other files don't make them different. A conversation that opens with an image
on its own goes round robin until a turn adds text. If your clients open every
conversation the same way, send a conversation ID in a header and key on it with
an [expression](#sticky-by-expression) instead.

A Responses API request that continues a stored response with
`previous_response_id` or `conversation` sends only the new turn, so its key
changes on every turn. See [Stored responses](#stored-responses).

The policy hashes this text in memory. It never stores or logs it.

### Sticky by expression

`"stickyBy": "expression"` keys on one value that `expression` selects from the
request or its context. Expressions use the same syntax as
[Metering budget expressions](/docs/policies/ai-gateway-metering-inbound#budget-expressions),
so a value you budget on can also keep requests on one model. For example, to
keep each end user on one model when your app sends an end-user ID in a header:

```json
{
  "strategy": "sticky",
  "stickyBy": "expression",
  "expression": "request.headers.get(\"x-end-user-id\")",
  "models": {
    "completions": ["azure-eastus/gpt-5", "azure-westus/gpt-5"]
  }
}
```

In JSON, escape the double quotes inside an expression, as the example does.
Common expressions:

| Expression                      | Selects                                                   |
| ------------------------------- | --------------------------------------------------------- |
| `request.headers.get("<name>")` | One request header, such as an end-user ID your app sends |
| `request.user.data.<property>`  | A property of the caller's metadata                       |
| `request.params.<name>`         | A path parameter from the matched route                   |
| `context.custom.<property>`     | A value your own policy puts on `context.custom`          |

See
[Supported expressions](/docs/policies/ai-gateway-metering-inbound#supported-expressions)
in the Metering docs for the full list. An expression is a selector, not code,
so it can't compare values, call string methods, or read the request body.

- The value follows Metering's value rules. It must be a string or a safe
  integer, and `-0` doesn't count. The gateway converts a string to well-formed
  NFC Unicode, then rejects it when it's longer than 256 UTF-8 bytes or contains
  a control character, U+2028, or U+2029.
- Values are case-sensitive, so `Acme` and `acme` can land on different models.
  Two strings that normalize to the same NFC text share a model.
- When the expression resolves to nothing, to an empty string, or to a value
  these rules reject, such as an object, a fractional number, or an overlong
  string, the request goes round robin. A stored-response request fails with
  `400` instead when the models that store responses have more than one provider
  name. The request log says why.
- Per-request identifiers, such as `context.requestId`, aren't allowed, because
  a key that never repeats can't keep requests together.
- An expression outside the syntax is an options error, and the error shows
  where parsing stopped.

The policy hashes the value in memory. It never stores or logs it.

## Failures and retries

Each request gets a retry model, except stored-response requests,
`/v1/messages/count_tokens` requests, `/v1/systemone` requests, and requests
that only one model in the list can serve. The AI Gateway treats the retry model
like any other backup model. When it moves there depends on the endpoint:

| When the chosen model                                        | Chat completions and embeddings                 | Responses and Messages                          |
| ------------------------------------------------------------ | ----------------------------------------------- | ----------------------------------------------- |
| Returns `429`                                                | Retries it twice, then moves to the retry model | Retries it twice, then moves to the retry model |
| Returns a `5xx`, `408`, or `425` status, or can't be reached | Moves to the retry model                        | Fails the request                               |
| Doesn't start responding within `fallbackTimeoutSeconds`     | Moves to the retry model                        | Keeps waiting                                   |
| Returns any other error, such as `400` or `401`              | Fails the request                               | Fails the request                               |

Responses and Messages requests use the provider's native API. On those
endpoints the AI Gateway moves to another model only for rate limits, as it does
without this policy.

- Each request tries at most two models. The gateway retries a rate-limited
  model twice, after 250 ms and then 500 ms. If every attempt returns `429`, the
  request makes six calls, three to each model, before the gateway returns
  `429`.
- Once a model starts streaming a response, the request stays on that model.
- A request that doesn't stream gets its response only when the model finishes,
  so on chat completions and embeddings `fallbackTimeoutSeconds` limits the
  whole response. A model that's still working when it runs out moves the
  request to the retry model, which starts over with twice the time, and the
  request fails if that runs out too. Raise `fallbackTimeoutSeconds` for slow
  models, such as reasoning models, or stream the response. When Fallback Model
  runs after this policy with a `fallback`, raise its `fallbackTimeoutSeconds`
  too. Its timeout replaces this policy's on requests that get its `fallback`.
  See [With Fallback Model](#with-fallback-model).
- When only one model in the list can serve a request, this policy sets no retry
  model and no timeout. On `/v1/responses` and `/v1/messages`, a list that mixes
  providers or vendors often has only one, so list at least two models for each
  API your clients use. Fallback Model can also add a backup. See
  [With Fallback Model](#with-fallback-model).
- The policy doesn't remember failures between requests. While a model is down,
  every request that picks it tries it first: chat completions and embeddings
  requests then move to the retry model, and Responses and Messages requests
  fail. With `sticky` routing, the same callers pick the failing model every
  time.

## Endpoints

| Endpoint                                                                       | Behavior                                                                                                                                                                                                                                                              |
| ------------------------------------------------------------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `/v1/chat/completions`                                                         | Models the AI Gateway sends Chat Completions requests to are eligible, and the gateway translates the request for each provider. System One models are skipped. Claude on Vertex AI isn't skipped, and requests that pick it fail with `400`.                         |
| `/v1/responses`                                                                | Only models the AI Gateway sends Responses API requests to are eligible: OpenAI models, Azure AI and OpenRouter models other than Claude, and Bedrock Mantle models other than Claude. Other providers are skipped, even when their own API serves the Responses API. |
| `/v1/messages`, `/v1/messages/count_tokens`                                    | Only models the AI Gateway sends Anthropic Messages API requests to are eligible: Anthropic models, and Claude on Azure AI, Bedrock Mantle, Vertex AI, and OpenRouter. `count_tokens` also skips providers that don't count tokens, such as OpenRouter.               |
| `/v1/systemone`                                                                | Only TypeSafe System One (Jev) models are eligible, so you can balance several Jev models or TypeSafe accounts. Jev through OpenRouter isn't eligible. These requests get no retry model.                                                                             |
| `/v1/embeddings`                                                               | Uses the `embeddings` list.                                                                                                                                                                                                                                           |
| `GET` or `DELETE /v1/responses/{id}`, and `GET /v1/responses/{id}/input_items` | Routed to the account that holds the response, when the policy can tell which one it is. See [Stored responses](#stored-responses).                                                                                                                                   |
| `GET /v1/models`                                                               | Passes through unchanged.                                                                                                                                                                                                                                             |
| Bedrock Runtime (`/model/{modelId}/…`)                                         | Not supported. An app whose policy chain includes this policy rejects Bedrock Runtime requests. A model on a Bedrock Runtime provider can't serve any `/v1` endpoint, so the policy skips it.                                                                         |

If no model in the list can serve the request's endpoint, the policy returns
`400`. The message names the endpoint, and for an endpoint that uses a
provider's own API, such as `/v1/messages`, the API it requires.

These rules depend on the provider, and on whether the model is Claude. The
policy doesn't check which APIs each model serves. A request that picks a model
on an API the model doesn't serve fails, such as a `/v1/chat/completions`
request that picks a model that serves only the Responses API.

On `/v1/responses`:

- Bedrock Mantle serves the Responses API for only some of its models other than
  Claude. Requests that pick a Mantle model without Responses support fail. To
  check a model, read the **APIs supported** for the `bedrock-mantle` endpoint
  on its
  [AWS model card](https://docs.aws.amazon.com/bedrock/latest/userguide/model-cards.html).
  The AI Gateway calls that endpoint, and the `bedrock-runtime` endpoint's list
  can differ.
- OpenRouter stores no responses. When the list also has a model that does,
  OpenRouter gets only creates that set `store` to `false`. See
  [Stored responses](#stored-responses).

### Stored responses

A stored response lives on the provider account that created it. Requests that
read or continue one must reach that account: retrieving a response, listing its
input items, deleting it, and creating a response with `previous_response_id`, a
`conversation`, or an `item_reference` input item.

| When                                                                        | Stored-response requests                                                  |
| --------------------------------------------------------------------------- | ------------------------------------------------------------------------- |
| Every model in the list that stores responses has the same provider name    | Go to that provider                                                       |
| `strategy` is `sticky` by `user` or `expression`, and the request has a key | Go to the model the key picks, where that caller's responses were created |
| Anything else                                                               | Fail with `400`, and the request log names the fix                        |

- The provider name is the part before the slash, compared without regard to
  case. `openai/gpt-5` and `openai/gpt-5-mini` share one, but
  `azure-eastus/gpt-5` and `azure-westus/gpt-5` are two.
- Stored-response requests get no retry model, because no other account holds
  the response. Fallback Model placed after this policy still adds its
  `fallback`, so a continuation that stays rate limited moves to the fallback,
  and fails there if the fallback is on another account.
- A provider that stores nothing, such as OpenRouter, never gets stored-response
  requests. It gets a create only when the create sets `store` to `false`, or
  when no provider in the list stores responses, because a create stores its
  response unless it opts out.
- A create that moves to the retry model stores its response on the retry
  model's account. If that's another account, later requests for the response go
  to the chosen model's account and fail.
- The retry model, a `fallback`, or a `quotaFallback` is on another account when
  its provider name differs from the chosen model's. It's also on another
  account when AI Gateway Authentication passes callers' own API keys through,
  because the chosen model uses the caller's key and the others use your
  configured key.
- The `400` tells the caller only that this API can't read or continue the
  stored response. The fix is yours to make, so the request log carries it.
- `conversation` keys can't route them to one of several provider names. A
  retrieve sends no messages, and a request with `previous_response_id` sends
  only the new turn. With one provider name, each continuation is keyed on its
  new turn, so it can run on a different model than the turn before.
- Changing the list can move a caller to another model. Responses stored before
  the change stay on the old account.

## Policy order

Place Model Load Balancing after authentication and before the policies that
read or change model routing:

```json
{
  "inboundPolicyChain": [
    { "name": "ai-gateway-auth-inbound" },
    { "name": "ai-gateway-load-balancing-inbound" },
    { "name": "ai-gateway-fallback-model-inbound" },
    { "name": "ai-gateway-metering-inbound" },
    { "name": "ai-gateway-semantic-cache-inbound" }
  ]
}
```

This chain shows the order only. Entries without `options` use the options
declared in `config/policies.json`. See
[Configure lists per app](#configure-lists-per-app).

| Policy                                                               | With Model Load Balancing                                                                                                                                                                                                                                                                                                          |
| -------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| [Authentication](/docs/policies/ai-gateway-auth-inbound)             | Place it before, so `request.user.sub` is set for sticky routing by user.                                                                                                                                                                                                                                                          |
| [Model Override](/docs/policies/ai-gateway-model-override-inbound)   | Not needed, because this policy picks the model for every request it handles. If Model Override runs first, its model wins: always with `force`, and with `default` when a request omits `model`, as stored-response reads do. Placed after, it does nothing for request types this policy has a list for.                         |
| [Model Filtering](/docs/policies/ai-gateway-model-filtering-inbound) | Not needed for request types this policy has a list for, because the list already limits the models clients can reach. If Model Filtering runs first, it routes or refuses each request by the model the client asked for, and this policy does nothing. Placed after, it doesn't check this policy's choices.                     |
| [Fallback Model](/docs/policies/ai-gateway-fallback-model-inbound)   | Place it after, where it changes the retry model, the timeout, and where over-quota requests go. Placed before, it does nothing. See [With Fallback Model](#with-fallback-model).                                                                                                                                                  |
| [Smart Router](/docs/policies/ai-gateway-smart-router-inbound)       | When it routes a request, its model wins, whichever runs first. The request loses this policy's retry model, and Smart Router doesn't apply this policy's endpoint or stored-response rules. Fallback Model placed after Smart Router still adds its `fallback` and `quotaFallback`. Placed before Smart Router, they're lost too. |
| [Metering](/docs/policies/ai-gateway-metering-inbound)               | Place it after this policy and Fallback Model, so a request over its quota can switch to Fallback Model's `quotaFallback`. Costs and usage use the model that served the request.                                                                                                                                                  |
| [Semantic Cache](/docs/policies/ai-gateway-semantic-cache-inbound)   | Place it after. Its cache key uses the request's own `model`, not this policy's choice, so with a list of different models a cached answer can come from any of them. Placed before, it serves a cache hit without this policy when the request's `model` is a valid model, and skips the hit otherwise.                           |

### With Fallback Model

Fallback Model placed after this policy changes where failed and over-quota
requests go:

| Requests                                                                                          | Fallback Model's `fallback`                                                                       | Its `quotaFallback`, over quota                                                                                                                 |
| ------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------- |
| Chat completions and embeddings                                                                   | Becomes the retry model, and its `fallbackTimeoutSeconds`, 60 when unset, replaces this policy's. | Gets the request. A `quotaFallback` that the AI Gateway refuses on chat completions, such as Claude on Vertex AI, fails the request with `400`. |
| Responses and Messages creates                                                                    | Becomes the retry model if it can serve the endpoint. Otherwise the request has no retry model.   | Gets the request if it can serve the endpoint. Otherwise the gateway returns `429`.                                                             |
| `/v1/messages/count_tokens`                                                                       | Never tried.                                                                                      | Gets the request if it can serve the endpoint. Otherwise the gateway returns `429`.                                                             |
| `/v1/systemone`, and requests that retrieve, delete, or list the input items of a stored response | Never tried.                                                                                      | Never used. The gateway returns `429`.                                                                                                          |

- Requests this policy sends to the `fallback` model itself keep this policy's
  retry model and timeout. Requests it sends to the `quotaFallback` model itself
  get no quota fallback, so when they're over quota, the gateway returns `429`.
  Use a `quotaFallback` that isn't in the list.
- When a `quotaFallback` call fails in a way the endpoint retries, the request
  moves to the retry model, so an over-quota request can still run on a model in
  the list. To keep over-quota requests off the list, give Fallback Model a
  `fallback` that isn't in the list. To keep failed requests inside the list,
  give Fallback Model only a `quotaFallback`.
- No model serves both `/v1/responses` and `/v1/messages`, so when your clients
  call both, a `fallback` or a `quotaFallback` covers at most one of them.
- Over quota, a create that continues a stored response goes to the
  `quotaFallback`, and fails there if the `quotaFallback` is on another account.
  Other creates over quota go there too, so a response they store lives on the
  `quotaFallback`'s account. If that's another account, requests for the
  response fail once they're no longer over quota. See
  [Stored responses](#stored-responses).
- For embeddings, set `fallback` to the same embedding model as the list. With a
  one-model list, use that model on another provider name, because a `fallback`
  equal to the chosen model is skipped.
- Unless you've changed them, Fallback Model's declared options set
  `openai/gpt-4o-mini` as the `completions` `fallback`. A Fallback Model entry
  without `options`, like the one in the chain above, makes that model the retry
  model. On a gateway that can't use `openai/gpt-4o-mini`, such as one with no
  provider named `openai`, every request that Fallback Model adds it to fails
  with `500`, even requests that never try a backup. Give the entry its own
  `options`.

## Configure lists per app

Declare the policy in `config/policies.json` with a default list. An app can
supply its own options in its chain entry. Entry options replace the declared
options entirely:

```json
{
  "inboundPolicyChain": [
    {
      "name": "ai-gateway-load-balancing-inbound",
      "options": {
        "strategy": "sticky",
        "stickyBy": "conversation",
        "models": {
          "completions": ["azure-eastus/gpt-5", "azure-westus/gpt-5"]
        }
      }
    }
  ]
}
```

Unless you've changed them, the declared options are a one-model list,
`openai/gpt-4o-mini`. A chain entry without `options` then sends every chat
completions and Responses request to that model, and fails Messages and System
One requests with `400`, because that model can't serve them.

The gateway validates chain options when the policy runs. Invalid options fail
requests with an error that names the field to fix. So does a model the gateway
can't use, such as an unknown provider or an inactive model, on every request
that picks it as the chosen or the retry model. A provider whose API key isn't
on the gateway yet fails those requests with an error that says the provider
isn't ready, and the request log names the provider and its environment
variable.

## Monitoring

AI Gateway analytics attribute each request to the model that served it. Logging
plugins also receive these properties on the request's log entries after the
policy picks a model:

| Property                    | Value                              |
| --------------------------- | ---------------------------------- |
| `aiLoadBalancingModel`      | The model the policy chose         |
| `aiLoadBalancingRetryModel` | The retry model, when there is one |

The properties record this policy's picks. They don't change when Fallback Model
replaces the retry model or Smart Router replaces the choice.

## Limitations

- Each list can have up to 32 models, each listed once. The policy compares
  entries without regard to case, so `openai/gpt-5` and `OpenAI/gpt-5` count as
  one model. Provider names work in any case, but write each model ID exactly as
  your provider spells it. A model ID in another case, such as `openai/GPT-5`,
  can fail to route.
- Each request tries at most two models.
- An expression can be up to 1024 bytes and 32 property segments.

Lists that mix providers or models have limits of their own:

| When                                                                                     | What happens                                                                                                                                                                                                                                                                                                                                                                                  |
| ---------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| A list mixes providers or vendors, on `/v1/responses` or `/v1/messages`                  | Only models the AI Gateway sends that API's requests to are eligible, and that depends on the model, not only its provider. On Azure AI, for example, Claude models are eligible on `/v1/messages`, and GPT models on `/v1/responses`. See [Endpoints](#endpoints). A mixed list often leaves one model and no retry model, so list at least two models for each API your clients use.        |
| A list mixes vendors, on `/v1/chat/completions`                                          | Each pick goes through its provider's translation, which drops what that provider doesn't support. Claude models ignore `response_format` and `seed`, use a `max_tokens` of 1,024 unless the request sets `max_tokens`, read `developer` messages as user messages, and fail requests that carry images, audio, or files. For the same behavior on every pick, list deployments of one model. |
| The models accept different parameters or context lengths                                | A request that one model rejects fails whenever the policy picks that model. The retry model isn't tried.                                                                                                                                                                                                                                                                                     |
| AI Gateway Authentication passes callers' own API keys through (`credentialPassthrough`) | The caller's key goes to whichever model the policy picks, so list only models that accept it. Every model that can be the retry model also needs its own configured key, and requests that move there run on your provider account.                                                                                                                                                          |
| You want different lists for Chat Completions, Responses, and Messages                   | One `completions` list serves them all. A second Model Load Balancing policy in the chain doesn't run for request types the first one has a list for.                                                                                                                                                                                                                                         |

Read more about [how policies work](/articles/policies)
