Fallback Models
Fallbacks let an app keep serving requests when its primary model fails, times out, or an applicable usage limit is exceeded, instead of returning an error to the caller. They're configured with the Fallback Model policy in the app's policy chain.
The AI Gateway offers two independent fallback mechanisms, each triggered by a different condition:
| Mechanism | Triggers when… | Without a fallback set… |
|---|---|---|
| Fallback & Timeout | The primary fails with a retryable error (5xx, 408, 425, 429), a network failure, or a timeout | The gateway returns the error to the caller |
| Quota Fallback | An app, team, or gateway usage limit is exceeded | The request is blocked with a 429 |
Both are configured entirely in the Zuplo Portal, and either can route to any
provider—the fallback doesn't have to share the primary's provider. Fallback
models are referenced like any model, as providerName/model.
Configure fallbacks
-
Open the Apps & Teams tab of your AI Gateway project and select the app to edit.
-
Select the Policies tab. If the chain doesn't have a Fallback Model policy yet, click Add Policy and add it—placed directly after Model Filtering.
-
Configure the policy:
- Fallback: the model attempted after a retryable error or timeout, for
example
anthropic/claude-haiku-4-5. - Quota fallback: the model used once a usage limit is exceeded, for
example
openai/gpt-4o-mini. - Fallback timeout (seconds): how long the primary call can run before
the gateway fails over. The default is
60; values from1to300are accepted.
- Fallback: the model attempted after a retryable error or timeout, for
example
-
Save. The change applies within about a minute.
The request timeout applies only when a fallback model is set. If no fallback is configured, the primary model call runs unbounded.
Quota fallback
For requests continuing to a provider, the quota fallback selects an alternate,
usually cheaper, model when an app, team, or gateway
usage limit is exceeded, rather than blocking the request
with a 429. This keeps an app available after an applicable budget, token, or
request threshold is crossed, while shifting the overflow traffic to a
lower-cost model.
The Fallback Model policy supplies the quota fallback selection. Add
Budgets and Costs to set the app's own limits, and place it after Fallback
Model. If you leave the quota fallback empty, the gateway blocks requests with a
429 once an applicable limit is exceeded.
If a Block budget is exhausted, cache hits return 429 instead of the
cached answer. Quota fallback doesn't apply to cache hits. See
Successful cache responses.
Point the quota fallback at a smaller, cheaper model so overflow traffic stays inexpensive while remaining available. The fallback's own usage still counts toward every applicable limit.
How the two fallbacks combine
The mechanisms are evaluated independently and can both be active on the same app:
- A request that's over quota routes to the quota fallback model.
- A request that's within quota but hits an error or timeout on the primary routes to the error and timeout fallback model.
Set whichever fallbacks match the failure modes you want to protect against. Neither is required.
Error and timeout fallback applies to Chat Completions and Embeddings requests.
Requests to the native /v1/messages and /v1/responses endpoints pass through
without retrying a backup model; the quota fallback still applies to them.
Related resources
- Policy Chains - How the app's policy chain executes and the recommended policy order.
- Usage Limits - Configure the budget, token, and request limits that trigger a quota fallback.
- Custom Providers - Add your own provider to use as a primary or fallback model.