Fallback Models
Fallbacks let an app keep serving requests when its primary model fails, times out, or an applicable usage limit is exceeded, instead of returning an error to the caller. They're configured with the Fallback Model policy in the app's policy chain.
The AI Gateway offers two independent fallback mechanisms, each triggered by a different condition:
| Mechanism | Triggers when… | Without a fallback set… |
|---|---|---|
| Fallback & Timeout | The primary fails with a retryable error (5xx, 408, 425, 429), a network failure, or a timeout | The gateway returns the error to the caller |
| Quota Fallback | An app, team, or gateway usage limit is exceeded | The request is blocked with a 429 |
Both are configured entirely in the Zuplo Portal, and either can route to any
provider—the fallback doesn't have to share the primary's provider. Fallback
models are referenced like any model, as providerName/model.
Configure fallbacks
-
Open the Apps & Teams tab of your AI Gateway project and select the app to edit.
-
Select the Policies tab. If the chain doesn't have a Fallback Model policy yet, click Add Policy and add it—placed directly after Model Filtering.
-
Configure the policy:
- Fallback: the model attempted after a retryable error or timeout, for
example
anthropic/claude-haiku-4-5. - Quota fallback: the model used once a usage limit is exceeded, for
example
openai/gpt-4o-mini. - Fallback timeout (seconds): how long the primary call can run before
the gateway fails over. The default is
60; values from1to300are accepted.
- Fallback: the model attempted after a retryable error or timeout, for
example
-
Save. The change applies within about a minute.
The request timeout applies only when a fallback model is set. If no fallback is configured, the primary model call runs unbounded.
Quota fallback
The quota fallback routes requests to an alternate, usually cheaper, model when
an app, team, or gateway usage limit is exceeded, rather
than blocking the request with a 429. This keeps an app available after an
applicable budget, token, or request threshold is crossed, while shifting the
overflow traffic to a lower-cost model.
The Fallback Model policy supplies the quota fallback selection. Team and
gateway limits can activate it whether or not the chain includes Budgets and
Costs. Add Budgets and Costs when the app also needs its own limits, and place
it after Fallback Model. If you leave the quota fallback empty, the gateway
blocks requests with a 429 once an applicable limit is exceeded.
Point the quota fallback at a smaller, cheaper model so overflow traffic stays inexpensive while remaining available. The fallback's own usage still counts toward every applicable limit.
How the two fallbacks combine
The mechanisms are evaluated independently and can both be active on the same app:
- A request that's over quota routes to the quota fallback model.
- A request that's within quota but hits an error or timeout on the primary routes to the error and timeout fallback model.
Set whichever fallbacks match the failure modes you want to protect against. Neither is required.
Error and timeout fallback applies to Chat Completions and Embeddings requests.
Requests to the native /v1/messages and /v1/responses endpoints pass through
without retrying a backup model; the quota fallback still applies to them.
Related resources
- Policy Chains - How the app's policy chain executes and the recommended policy order.
- Usage Limits - Configure the budget, token, and request limits that trigger a quota fallback.
- Custom Providers - Add your own provider to use as a primary or fallback model.