Smart Router Policy
AI Gateway Policy
This policy is for use with the AI Gateway. See the AI Gateway documentation to learn how to configure and govern AI models with Zuplo.
Classifies the last user prompt with an AI Gateway app, an OpenAI-compatible
Chat Completions API, or Jev from TypeSafe, stores the result on
AIGatewaySmartRouter for later policies, and optionally routes completions by
classified complexity.
Configuration
The configuration shows how to configure the policy in the 'policies.json' document.
Code
Policy Configuration
name<string>- The name of your policy instance. This is used as a reference in your routes.policyType<string>- The identifier of the policy. This is used by the Zuplo UI. Value should beai-gateway-smart-router-inbound.handler.export<string>- The name of the exported type. Value should beAIGatewaySmartRouterInboundPolicy.handler.module<string>- The module containing the policy. Value should be$import(@zuplo/runtime).handler.options<object>- The options for this policy. See Policy Options below.
Policy Options
The options for this policy are specified below. All properties are optional unless specifically marked as required.
classifier<object>- Where the prompt is classified. Set exactly one ofapp,chatCompletions, orjev.app.adapterselects the API that classifier app exposes. Setting more than one section is a configuration error.app<object>- Classify with another AI Gateway app in this project. The call runs in-process and never leaves the gateway, so the classifier uses that app's providers, routing, and quotas.appId(required)<string>- The id of the AI Gateway app that runs the classifier. It must not be this app, and that app must not run this policy.model(required)<string>- The model (providerName/model) the classifier app uses to evaluate the user's request.providerNameis the name of the provider as configured on the classifier app, not a fixed vendor id. E.g.openai/gpt-4o-mini, orjev/jev-latestwhen the Jev provider is namedjev.apiKey<string>- API key sent asAuthorization: Bearerwhen invoking the classifier app. Omit it when the classifier app runs the Internal Only policy: the classifier call never leaves the gateway, so that policy accepts it and rejects everything arriving over the network, and no credential is sent. Set it only when the classifier app authenticates callers with API keys.adapter<string>- API contract exposed by the classifier app. chatCompletions calls /v1/chat/completions; systemOne calls /v1/systemone for models such as Jev. The app owns provider credentials, pricing, metering and analytics. Allowed values arechatCompletions,systemOne. Defaults to"chatCompletions".
chatCompletions<object>- Deprecated: useclassifier.appinstead. Classify by calling an OpenAI-compatible Chat Completions API directly, outside the gateway. The model must support structured outputs (response_formatwith a JSON schema).apiKey(required)<string>- API key for the service atbaseUrl, sent asAuthorization: Bearer.model(required)<string>- The model id the service expects, such asgpt-4o-mini. This is sent to the service as-is, not as aproviderName/modelreference.baseUrl<string>- Base URL of the OpenAI-compatible API. The classifier request is sent to{baseUrl}/chat/completions. Defaults to"https://api.openai.com/v1".
jev<object>- Deprecated: useclassifier.appwithadaptersystemOneinstead. Classify with Jev, TypeSafe's classification model, by calling a System One API directly: TypeSafe's by default, or another host such as OpenRouter throughbaseUrl. Jev answers typed questions instead of generating text, so it takes noclassifierPrompt.apiKey(required)<string>- API key for the service atbaseUrl, sent asAuthorization: Bearer: a TypeSafe key by default, or an OpenRouter key whenbaseUrlis OpenRouter.model<string>- The Jev model to call.jev-latestfollows the newest stable release on both TypeSafe and OpenRouter. Pin a versioned id to keep answers stable after tuningminConfidenceForRouting:jev-1.13.0on TypeSafe, orjev-1.13on OpenRouter. Defaults to"jev-latest".baseUrl<string>- API root of the System One API that serves Jev, the same value the TypeSafe SDK'sbaseURLtakes. The classifier request is sent to{baseUrl}/v1/systemone. Set it tohttps://openrouter.ai/apito call Jev through OpenRouter. Defaults to"https://api.typesafe.ai".
modelsByComplexity(required)<object>- Enables which model will be used based on the classifier results (low,mediumandhigh). This configuration is applied only whensmartRoutingEnabledis true and confidence score meets theminConfidenceForRoutingthreshold.low(required)<string>- Model used for prompts classified as low complexity.medium(required)<string>- Model used for prompts classified as medium complexity.high(required)<string>- Model used for prompts classified as high complexity.
smartRoutingEnabled<boolean>- When true, apply model routing based on the models set inmodelsByComplexitywhen confidence score meets theminConfidenceForRoutingthreshold. When set to false, classification still runs but model routing is not applied, useful for debugging or testing the classifier prompt. Defaults totrue.intents<object[]>- Dictionary of intents used to classify the user message being evaluated. Omit to use the built-in dictionary (code, summarization, translation, qa, conversation, classification, creative_writing, agentic, document_qa, other).id(required)<string>- Label used to classify the intent of the user message being evaluated.description(required)<string>- Short description of the intent label.
classifierPrompt<undefined>- Prompt for classifier.app (its default or chatCompletions adapter) and classifier.chatCompletions. Omit to use the built-in prompt. classifier.app.adapter systemOne and classifier.jev take no prompt; combining them with classifierPrompt is a configuration error.minConfidenceForRouting<number>- Minimum confidence (0–1) required before applying model routing. Unknown intents are capped strictly below this threshold. Defaults to0.5.classifierTimeoutMs<integer>- Timeout configured for the classifier task. When exceeded, the request is forwarded without classification to the original model. Defaults to8000.maxPromptChars<integer>- Maximum characters of user message sent to the classifier. Longer messages get truncated. Defaults to8000.classifierAppID<string>- **Deprecated**: useclassifier.app.appIdinstead. Still honored whenclassifieris not set; setting both is a configuration error.classifierModel<string>- **Deprecated**: useclassifier.app.modelinstead. Still honored whenclassifieris not set; setting both is a configuration error.classifierAppApiKey<string>- **Deprecated**: useclassifier.app.apiKeyinstead. Still honored whenclassifieris not set; setting both is a configuration error.
Using the Policy
Use this policy to classify the last user prompt on Chat Completions, Responses, and Anthropic Messages requests. Using the classification results, configure where to route the request based on its complexity.
Smart routing is on by default: the policy overwrites completions routing based
on the configuration in modelsByComplexity. Set smartRoutingEnabled to
false to keep classifying without routing, which is useful when testing a
classifier prompt.
Classification failure handling. If the message classification fails or times out, the request is forwarded to the original model.
Required options
classifier— where the prompt is classified. Set exactly one ofapp,chatCompletions, orjev; see Choose a classifier.modelsByComplexity—providerName/modelfor each oflow,medium, andhigh. Used for routing unlesssmartRoutingEnabledis set tofalse.
Omit intents and classifierPrompt to use the built-in dictionary (code,
summarization, translation, qa, conversation, classification, creative_writing,
agentic, document_qa, other) and the built-in classification prompt.
Choose a classifier
| Classifier | Set | Use it when |
|---|---|---|
classifier.app | appId, model | You want the classifier to use one of your providers, through an AI Gateway app in this project. The call stays in-process. Set adapter: "systemOne" for Jev and other System One models. |
classifier.chatCompletions (deprecated) | apiKey, model | You need to call an OpenAI-compatible endpoint directly with its own key, without a second AI Gateway app. |
classifier.jev (deprecated) | apiKey | You need to call Jev, from TypeSafe, directly with its own key. TypeSafe and OpenRouter both serve it. |
The three are mutually exclusive. Setting more than one is rejected with a
ConfigurationError rather than guessed at, because routing on a classifier you
didn't mean to use would look like it was working.
Use classifier.app for new configurations. The direct chatCompletions and
jev sections are deprecated: they bypass the classifier app's policy chain,
budgets, and analytics, and new classifier features are added only to
classifier.app.
AI Gateway app
Code
appId— the AI Gateway app that runs the classifier.adapter— optional:chatCompletions(default) orsystemOne. Selects/v1/chat/completionsor/v1/systemoneon that app.model— theproviderName/modelthat app routes the classifier request to.apiKey— optional. See Authenticating the classifier app.
Give the classifier its own application rather than pointing this at the app the policy runs on. The classifier request never leaves the gateway, so that application needs no API key checked at all — put the Ensure Gateway Internal Invocation Only policy on it in place of API Key Authentication.
For Jev registered with your gateway, use the same app path with System One:
Code
Replace jev with your configured provider name. The classifier app resolves
the provider assignment and owns upstream credentials, routing, DB pricing,
budget checks, token metering, and analytics. No direct provider key, base URL,
or pricing configuration is needed in Smart Router. Its result still includes
the classifier's usage and cost; it does not emit duplicate analytics under the
calling app.
Smart Router sends its intent and complexity questions to the app and interprets
the returned System One answers. Omit classifierPrompt for this adapter. Keep
Smart Router off the classifier app's own chain. System One is not a chat-shaped
request, so choose policies compatible with that endpoint; existing guardrails
that cannot inspect its shape still reject it according to their configuration.
Prefer this app configuration for registered gateway providers. The direct sections below are deprecated and remain only for existing configurations.
Chat Completions API (deprecated)
Code
apiKey— sent asAuthorization: Bearer.model— the model id the service expects, such asgpt-4o-mini, not aproviderName/modelreference. It must support structured outputs: the classifier request asks for a strict JSON schema inresponse_format.baseUrl— optional,https://api.openai.com/v1by default. Set it to call another OpenAI-compatible service. The request goes to{baseUrl}/chat/completions.
To classify with an app in this project, use classifier.app rather than
pointing baseUrl at the app's URL. classifier.app invokes the app
in-process, and it rejects a classifier app that would call back into this
policy.
Jev (deprecated)
Code
apiKey— your TypeSafe API key, from the TypeSafe dashboard, sent asAuthorization: Bearer. Through OpenRouter, your OpenRouter key.model— optional,jev-latestby default, which follows the newest stable release. After you tuneminConfidenceForRouting, pin a versioned model so a new release can't move answers under your threshold:jev-1.13.0on TypeSafe, orjev-1.13on OpenRouter.baseUrl— optional,https://api.typesafe.aiby default. It's the API root, the same value the TypeSafe SDK'sbaseURLtakes, and the request goes to{baseUrl}/v1/systemone.
To call Jev through OpenRouter, set baseUrl to OpenRouter's API root and use
an OpenRouter key. OpenRouter serves the same System One API and bills your
OpenRouter account:
Code
The policy sends the prompt and two questions to {baseUrl}/v1/systemone:
- Intent — a Choice over your intents. Each
idis an option, and itsdescriptiontells Jev what the option covers. - Complexity — a Score over three levels, low to high. The policy rounds the score to the nearest level.
profile.confidence is the lower of Jev's confidence in the two answers, so
routing applies only when Jev is sure of both. Jev writes no reasons, so
profile.reasons is empty.
Jev takes no system prompt, so setting classifierPrompt with classifier.jev
is rejected with a ConfigurationError. The same restriction applies to
classifier.app.adapter: "systemOne".
chatCompletions and jev call a service outside the gateway with the key you
configure, and send it the prompt. When that service answers with an error, the
policy logs the status and its most likely cause but not the response body,
which can echo the request back, prompt included.
To migrate a direct section, register the provider with the gateway, create a
classifier app for it, and replace the section with classifier.app. Use
adapter: "systemOne" to replace classifier.jev. Existing classifier.app
configurations continue to default to Chat Completions when adapter is
omitted. systemOne uses the existing Jev request and response contract; it
does not add support for arbitrary decision APIs.
Authenticating the classifier app
The classifier call is made with context.invokeRoute and never leaves the
gateway, so there is no network hop to protect. Give the classifier application
the Internal Only policy instead of AI
Gateway Authentication: it accepts requests this gateway made itself and rejects
everything that arrives over the network. Then omit classifier.app.apiKey —
nothing is sent, and there is no key to rotate or leak.
Set classifier.app.apiKey only when the classifier application authenticates
callers with API keys. It is sent as Authorization: Bearer, typically from
$env(CLASSIFIER_APP_API_KEY). Omit it, or leave it empty, and the classifier
request carries no Authorization header. If you leave it unset while the
classifier app still checks keys, that app answers 401 and the gateway logs
which policy to add.
classifier.app.appId must not be this app's own id, and the classifier
application must not run this policy itself. Both are rejected with a
ConfigurationError instead of silently skipping or failing open, so a
misconfigured routing loop — within one app, or between two — is never mistaken
for smart routing that is simply doing nothing.
The classifier application must not run
Akamai AI Firewall either.
That firewall would inspect the internal classification call as if it were a
completion request — scanning the classifier's system prompt and JSON-only
response for injection/malicious-content patterns — and a resulting denial makes
Smart Router fail open, indistinguishable from smart routing quietly doing
nothing. This is rejected with a ConfigurationError for the same reason.
Example
Code
Migrating from classifierAppID
Earlier versions configured the classifier app with three top-level options.
They're deprecated, and still honored when classifier isn't set:
| Deprecated option | Replacement |
|---|---|
classifierAppID | classifier.app.appId |
classifierModel | classifier.app.model |
classifierAppApiKey | classifier.app.apiKey |
Move all three at once. Setting classifier together with any of them is
rejected with a ConfigurationError.
Policy order
Code
Smart Router always set()s routing when smart routing applies, even if
filtering already selected a model. Filtering skips when routing is already set,
so putting this policy first would also skip allow-list checks on the client's
original model.
How classification drives routing
Complexity is classified independently of intent, as low, medium, or high.
Smart Router looks up modelsByComplexity[complexity] and applies it only when
all of the following hold:
smartRoutingEnabledistrue(the default).- The classified intent is a known one (not capped as an unknown intent).
modelsByComplexityhas a model configured for that complexity.profile.confidence >= minConfidenceForRouting(default0.5).
When routing isn't applied, read smartRouting.reason from the result (see
Using classification results in custom code)
to see why: disabled, unknown-intent, no-model, low-confidence,
internal-error, or applied.
Other advanced options
| Option | Default | Purpose |
|---|---|---|
smartRoutingEnabled | true | Set it to false to classify every request and record the result without changing which model serves the request. |
minConfidenceForRouting | 0.5 | Raise it to route only on confident classifications; lower it to route more aggressively. Unknown intents are always capped just below this value, so they never qualify regardless of the setting. |
classifierTimeoutMs | 8000 | How long to wait for the classifier before giving up and forwarding the request unclassified. |
maxPromptChars | 8000 | Truncates the prompt sent to the classifier, to keep classifier cost and latency bounded on very long prompts. |
Advanced configuration
Use these options to control what the classifier evaluates and how strongly its result influences routing.
Custom intents
intents replaces the built-in dictionary entirely — it's not additive. Provide
a non-empty list of { id, description } pairs:
Code
The id values become the enum the classifier model must return — the
classifier calls a strict JSON-schema chat completion, so it can only return one
of your configured ids. The description values are only shown to the
classifier if your prompt includes {{intents}} (see below).
With classifier.jev or classifier.app.adapter: "systemOne", the id values
are the options of Jev's intent question, and every description is shown to
Jev with its option. There is no prompt, so {{intents}} doesn't apply.
If the classifier returns an id outside this list, Smart Router keeps it as an
"unknown intent" and caps its confidence just below minConfidenceForRouting,
so it's never eligible for routing.
Custom classifier prompt
classifierPrompt replaces the built-in classifier prompt for classifier.app
(its default or chatCompletions adapter) and classifier.chatCompletions.
classifier.app.adapter: "systemOne" and classifier.jev take no prompt, so
combining classifierPrompt with either is a configuration error. It accepts
either a string or an array of lines:
Code
Include the literal placeholder {{intents}} anywhere in the prompt to have it
replaced with a - id: description line for each configured intent (or the
built-in ones, if intents is also omitted). Pairing a custom intents list
with the built-in classifierPrompt (by omitting classifierPrompt entirely)
works out of the box, because the built-in prompt already contains
{{intents}}.
Using classification results in custom code
AIGatewaySmartRouter.get(context) returns the result Smart Router stored on
the request, or undefined if it didn't run (non-AI request, unreadable body,
or a fail-open error):
Code
result has this shape:
Code
With classifier.jev or classifier.app.adapter: "systemOne",
classifierModel is the versioned model that answered, in the host's naming:
jev-1.13.0 on TypeSafe, or typesafe/jev-1.13-20260917 on OpenRouter. usage
reports the input and output tokens that the Jev API returns as inputTokens
and outputTokens.
Read it from any policy placed after Smart Router in the chain. Common uses:
-
Branch on intent or complexity — apply a stricter rate limit, a different DLP policy, or a longer timeout for
highcomplexity or anagenticintent:Code -
Layer business logic on top of
modelsByComplexity— for example, cap the model for free-tier callers regardless of classified complexity. Read the plan from the caller's API key metadata, not from a raw request header (the caller controls headers and could set or omit them to bypass the cap). The built-in API Key Auth policy puts a key's metadata onrequest.user.data— setplanthere when you create the key, and it lands on every request that key makes:Code -
Observe why routing wasn't applied — log when
smartRouting.reasonislow-confidenceorunknown-intentto tuneminConfidenceForRoutingor the intent taxonomy:Code -
Surface classification for debugging — add response headers in a non-production environment to see what the classifier returned. This runs in an outbound policy, so return a new
Responsecarrying the headers — mutating a clonedHeadersobject alone has no effect on what the caller receives:Code
What content is evaluated
The policy reads the last user message in order to classify its intent and complexity. Embeddings, tool messages and other non-AI paths are skipped.
Fail-open behavior
The policy never 500s the user request for an internal classifier problem. Invalid or incomplete options, classifier errors and timeouts, empty classifier responses, and smart routing catalog errors are logged and the original request continues. Chat Completions, Responses, and Anthropic Messages are classified automatically. Embeddings and other non-AI paths are skipped.
These misconfigurations throw a ConfigurationError instead, which fails the
request with a 500 that names the problem. Failing open on them would look
like smart routing quietly doing nothing:
- More than one section in
classifier(app,chatCompletions, andjev). - A deprecated
classifierAppID,classifierModel, orclassifierAppApiKeynext toclassifier. classifierPromptwithclassifier.app.adapter: "systemOne"orclassifier.jev.- A
classifier.app.appIdthat names this app. - A classifier app that runs this policy or Akamai AI Firewall.
One shape is refused rather than skipped. A Bedrock Runtime request
(/model/{modelId}/{operation}) carries a model-native body that the gateway
does not parse, so there is no prompt to classify, and letting it continue would
present your routing as in effect when it was not. The policy answers a
400 ValidationException in the AWS error envelope, naming the policy. Remove
the policy from that route to use Bedrock Runtime through it.
Classifier cost reporting
Use classifier.app for DB-backed pricing, metering, analytics and budgets.
After classification, AIGatewaySmartRouter.get(context) exposes token usage
and cost from the invoked gateway's accounting headers, including its accounted
input, output, and total tokens and cost accuracy. Smart Router does not
recompute usage from the provider response for app classifiers. Missing or
invalid token headers produce zero counts; missing cost headers produce
source: "unpriced". Responses from older gateways without cost accuracy retain
a conservative estimate for catalog costs.
Deprecated direct classifiers retain token usage and any validated
provider-reported cost in the Smart Router result, but do not look up DB prices,
send gateway analytics, or increment budget meters. If the provider does not
report a valid cost, the result is unpriced; its numeric zero does not mean
the call was free. No additional pricing configuration is supported for direct
calls.
App classifier results also include usage.breakdown when the gateway has
canonical accounting detail. It exposes uncached input, cache-read input,
cache-write input (with TTL when known), visible output, and reasoning output.
These disjoint categories are already included in inputTokens and
outputTokens; do not add them to the totals again. Unreported detail remains
absent, while an explicitly reported zero is preserved. Missing or malformed
breakdown headers do not discard aggregate token counts or cost. Smart Router
does not derive these categories from the provider body or infer cache TTLs.
Read more about how policies work