AI Gateway Model Override Policy
AI Gateway Policy
This policy is for use with the AI Gateway. See the AI Gateway documentation to learn how to configure and govern AI models with Zuplo.
AI Gateway Model Override pins the model an AI Gateway route uses. Set force
to send every request to one model no matter what the client selects, or
default to supply a model only when the request omits one. Place it before AI
Gateway Model Filtering and AI Gateway Fallback Model; a selection made by an
earlier policy is never overwritten.
Beta
This policy is in beta. You can use it today, but it may change in non-backward compatible ways before the final release.
Configuration
The configuration shows how to configure the policy in the 'policies.json' document.
Code
Policy Configuration
name<string>- The name of your policy instance. This is used as a reference in your routes.policyType<string>- The identifier of the policy. This is used by the Zuplo UI. Value should beai-gateway-model-override-inbound.handler.export<string>- The name of the exported type. Value should beAIGatewayModelOverrideInboundPolicy.handler.module<string>- The module containing the policy. Value should be$import(@zuplo/runtime).handler.options<object>- The options for this policy. See Policy Options below.
Policy Options
The options for this policy are specified below. All properties are optional unless specifically marked as required.
models(required)<object>- Override rules grouped by AI Gateway capability.completions<undefined>- Rule for chat completions, Responses, and Anthropic Messages requests.embeddings<undefined>- Rule for embedding requests.
Using the Policy
Use this policy when the gateway operator, not the client, decides which model an AI Gateway route uses. Each capability picks one mode:
forcesends every request to the configured model. A model in the request body is ignored.defaultsupplies the configured model only when the request omitsmodel. A request that selects a model keeps its selection.
providerName is the Provider Name configured in the Zuplo Portal. The text
after the first slash is the provider-specific model ID, so model IDs may
contain additional slashes.
Choosing a mode
Use force to:
- Pin a route to one vetted model regardless of client input.
- Repoint existing traffic at a new model without a client release.
- Expose a stable route like
/fast/v1/chat/completionswhose model you swap server-side.
Use default to:
- Let clients omit
modelwhile keeping full choice for clients that send one. - Give Responses management operations (
GET/DELETE /v1/responses/*), which have no request body, the routing they require.
Policy order
Place Model Override first among the model-selection policies:
Code
- A selection made by an earlier policy is never overwritten. When this policy sets the selection, AI Gateway Model Filtering leaves it unchanged, so a forced model does not also need an allow-list entry.
- In
defaultmode, a request that selects its own model passes through unchanged, and Model Filtering or the handler validates it as usual. The default never masks an invalid request model; a malformed value is still rejected downstream with a 400 response. - AI Gateway Fallback Model placed after this policy enriches the selection with
fallbackandquotaFallbackmodels. Keep fallback fields in that policy; this one accepts plainproviderName/modelstrings only. - A custom inbound policy placed before this one stays authoritative. For
example, an experiment policy can select routing for a fraction of traffic and
let
forcecatch the rest.
Options
models must contain completions, embeddings, or both. Each capability sets
exactly one of:
force- the model every request uses.default- the model used only when the request omitsmodel.
Values are plain providerName/model strings. Unsupported fields and malformed
references are rejected as configuration errors naming the field. If the policy
is attached but the route's capability has no rule, the policy passes the
request through unchanged.
Force example
Code
Default example
Code
Request behavior
| Situation | force | default |
|---|---|---|
Request omits model | Configured model is selected. | Configured model is selected. |
| Request selects a model | Configured model replaces it. | The request's model continues downstream. |
| Request model is malformed | Configured model is selected. | Passed through; rejected downstream with a 400. |
| Bodyless request (Responses management operations) | Configured model is selected. | Configured model is selected. |
| An earlier policy already selected routing | Skipped; earlier selection wins. | Skipped; earlier selection wins. |
| Route capability has no rule | Passed through unchanged. | Passed through unchanged. |
When the policy selects the configured model, the same routing validation used
everywhere else applies: the Provider Name must be configured, the model must be
available for the route's capability, and the Provider Assignment must have
usable credentials. A configured model that fails this validation is reported as
this policy's configuration error, naming the exact option such as
options.models.completions.force.
Native routes
/v1/responses requires a Provider Name backed by OpenAI, and /v1/messages
requires one backed by Anthropic. A configured model that selects an
incompatible provider type is a configuration error on every request it applies
to, so the policy reports it as one, naming the option to fix.
Observability
Responses report the model that actually served the request in their model
field, so a client can always see that an override applied. The policy also
writes a debug-level log entry with requestedModel and forcedModel when a
forced model replaces a request's differing selection.
Write your own override policy
Everything this policy does is built on the public
AIGatewayModelRouting.set(context, routing) primitive. Use a custom inbound
policy instead when the override depends on request data, for example routing by
header:
Code
Read more about how policies work