
# AI Gateway Model Override Policy

:::note{title="AI Gateway Policy"}

This policy is for use with the [AI Gateway](/docs/ai-gateway/overview). See
the AI Gateway documentation to learn how to configure and govern AI models
with Zuplo.

:::

AI Gateway Model Override pins the model an AI Gateway route uses. Set `force`
to send every request to one model no matter what the client selects, or
`default` to supply a model only when the request omits one. Place it before AI
Gateway Model Filtering and AI Gateway Fallback Model; a selection made by an
earlier policy is never overwritten.

:::caution{title="Beta"}

This policy is in beta. You can use it today, but it may change in non-backward compatible ways before the final release.

:::

## Configuration

The configuration shows how to configure the policy in the 'policies.json' document.

```json title="config/policies.json"
{
  "name": "my-ai-gateway-model-override-inbound-policy",
  "policyType": "ai-gateway-model-override-inbound",
  "handler": {
    "export": "AIGatewayModelOverrideInboundPolicy",
    "module": "$import(@zuplo/runtime)",
    "options": {
      "models": {}
    }
  }
}
```

### Policy Configuration

- `name` <code className="text-green-600">&lt;string&gt;</code> - The name of your policy instance. This is used as a reference in your routes.
- `policyType` <code className="text-green-600">&lt;string&gt;</code> - The identifier of the policy. This is used by the Zuplo UI. Value should be `ai-gateway-model-override-inbound`.
- `handler.export` <code className="text-green-600">&lt;string&gt;</code> - The name of the exported type. Value should be `AIGatewayModelOverrideInboundPolicy`.
- `handler.module` <code className="text-green-600">&lt;string&gt;</code> - The module containing the policy. Value should be `$import(@zuplo/runtime)`.
- `handler.options` <code className="text-green-600">&lt;object&gt;</code> - The options for this policy. [See Policy Options](#policy-options) below.

### Policy Options

The options for this policy are specified below. All properties are optional unless specifically marked as required.

- `models` **(required)** <code className="text-green-600">&lt;object&gt;</code> - Override rules grouped by AI Gateway capability.
  - `completions` <code className="text-green-600">&lt;undefined&gt;</code> - Rule for chat completions, Responses, and Anthropic Messages requests.
  - `embeddings` <code className="text-green-600">&lt;undefined&gt;</code> - Rule for embedding requests.

## Using the Policy

Use this policy when the gateway operator, not the client, decides which model
an AI Gateway route uses. Each capability picks one mode:

- `force` sends every request to the configured model. A model in the request
  body is ignored.
- `default` supplies the configured model only when the request omits `model`. A
  request that selects a model keeps its selection.

`providerName` is the Provider Name configured in the Zuplo Portal. The text
after the first slash is the provider-specific model ID, so model IDs may
contain additional slashes.

## Choosing a mode

Use `force` to:

- Pin a route to one vetted model regardless of client input.
- Repoint existing traffic at a new model without a client release.
- Expose a stable route like `/fast/v1/chat/completions` whose model you swap
  server-side.

Use `default` to:

- Let clients omit `model` while keeping full choice for clients that send one.
- Give Responses management operations (`GET`/`DELETE /v1/responses/*`), which
  have no request body, the routing they require.

## Policy order

Place Model Override first among the model-selection policies:

```text
Model Override -> Model Filtering -> Fallback Model -> AI Gateway handler
```

- A selection made by an earlier policy is never overwritten. When this policy
  sets the selection, AI Gateway Model Filtering leaves it unchanged, so a
  forced model does not also need an allow-list entry.
- In `default` mode, a request that selects its own model passes through
  unchanged, and Model Filtering or the handler validates it as usual. The
  default never masks an invalid request model; a malformed value is still
  rejected downstream with a 400 response.
- AI Gateway Fallback Model placed after this policy enriches the selection with
  `fallback` and `quotaFallback` models. Keep fallback fields in that policy;
  this one accepts plain `providerName/model` strings only.
- A custom inbound policy placed before this one stays authoritative. For
  example, an experiment policy can select routing for a fraction of traffic and
  let `force` catch the rest.

## Options

`models` must contain `completions`, `embeddings`, or both. Each capability sets
exactly one of:

- `force` - the model every request uses.
- `default` - the model used only when the request omits `model`.

Values are plain `providerName/model` strings. Unsupported fields and malformed
references are rejected as configuration errors naming the field. If the policy
is attached but the route's capability has no rule, the policy passes the
request through unchanged.

## Force example

```json
{
  "name": "ai-gateway-model-override-inbound",
  "policyType": "ai-gateway-model-override",
  "handler": {
    "export": "AIGatewayModelOverrideInboundPolicy",
    "module": "$import(@zuplo/runtime)",
    "options": {
      "models": {
        "completions": {
          "force": "openai/gpt-5"
        }
      }
    }
  }
}
```

## Default example

```json
{
  "models": {
    "completions": {
      "default": "anthropic/claude-haiku-4-5"
    },
    "embeddings": {
      "default": "openai/text-embedding-3-small"
    }
  }
}
```

## Request behavior

| Situation                                          | `force`                          | `default`                                       |
| -------------------------------------------------- | -------------------------------- | ----------------------------------------------- |
| Request omits `model`                              | Configured model is selected.    | Configured model is selected.                   |
| Request selects a model                            | Configured model replaces it.    | The request's model continues downstream.       |
| Request model is malformed                         | Configured model is selected.    | Passed through; rejected downstream with a 400. |
| Bodyless request (Responses management operations) | Configured model is selected.    | Configured model is selected.                   |
| An earlier policy already selected routing         | Skipped; earlier selection wins. | Skipped; earlier selection wins.                |
| Route capability has no rule                       | Passed through unchanged.        | Passed through unchanged.                       |

When the policy selects the configured model, the same routing validation used
everywhere else applies: the Provider Name must be configured, the model must be
available for the route's capability, and the Provider Assignment must have
usable credentials. A configured model that fails this validation is reported as
this policy's configuration error, naming the exact option such as
`options.models.completions.force`.

## Native routes

`/v1/responses` requires a Provider Name backed by OpenAI, and `/v1/messages`
requires one backed by Anthropic. A configured model that selects an
incompatible provider type is a configuration error on every request it applies
to, so the policy reports it as one, naming the option to fix.

## Observability

Responses report the model that actually served the request in their `model`
field, so a client can always see that an override applied. The policy also
writes a debug-level log entry with `requestedModel` and `forcedModel` when a
forced model replaces a request's differing selection.

## Write your own override policy

Everything this policy does is built on the public
`AIGatewayModelRouting.set(context, routing)` primitive. Use a custom inbound
policy instead when the override depends on request data, for example routing by
header:

```typescript
import {
  AIGatewayModelRouting,
  type ZuploContext,
  type ZuploRequest,
} from "@zuplo/runtime";

export default async function overrideModel(
  request: ZuploRequest,
  context: ZuploContext,
) {
  const tier = request.headers.get("x-plan-tier");
  await AIGatewayModelRouting.set(context, {
    completions: tier === "pro" ? "openai/gpt-5" : "openai/gpt-5-mini",
  });
  return request;
}
```

Read more about [how policies work](/articles/policies)
