Cookbooks

Cookbook: Point provider-native SDKs at the gateway

The AI Gateway serves every provider through the Universal API: OpenAI-shaped requests on /v1/chat/completions, Anthropic-shaped ones on /v1/messages, and so on. Four provider-native SDKs still refuse to talk to it—the AzureOpenAI class, @ai-sdk/azure with useDeploymentBasedUrls, @ai-sdk/google-vertex/anthropic, and @anthropic-ai/vertex-sdk—but not because the gateway can't handle their requests. It already speaks each client's body dialect. They fail only because each one hard-codes a provider-native URL grammar—a deployment in the path, a {model}:rawPredict suffix—that the gateway doesn't serve. A URL rewrite is all they need.

This recipe is a custom inbound policy that does that rewrite. Copy it into your own gateway and you own it: it's example code, not a runtime feature, and these clients aren't among the supported clients documented on the Azure AI or Vertex AI pages. Model references don't change—they stay the gateway's usual providerName/deployment for Azure (azureai/my-gpt) and providerName/publisher/model for Vertex (vertexai/anthropic/claude-haiku-4-5@20251001).

The policy

Copy this module into your gateway as modules/sdk-path-shim-inbound.ts. It rewrites the request path for the grammars it recognizes and leaves everything else untouched; the header comment lists every path it maps, and tests/sdk-shim.test.ts in zuplo/core is its executable half.

TypeScriptCode
import type { ZuploContext } from "@zuplo/runtime"; import { ZuploRequest } from "@zuplo/runtime"; /** * SDK path shim: rewrites the URL grammars of four vendor SDKs onto the AI * Gateway's own endpoints, so a client that hard-codes a provider-native path * can be pointed at the gateway by changing its endpoint and credential only. * * Copy this module into your gateway, register it in `policies.json` as a * `custom-code-inbound` policy, and list it FIRST in the catch-all route's * inbound policies — before `ai-gateway-configuration-loader-inbound`. The * loader answers 404 to any path whose segment after the app id is not `v1` * before it loads the app, so a shim placed in an app's stored chain would * never run for these paths. (This fixture mounts it on a sibling route, * `/sdk-shim/:app_id/(.*)`, only because `tests/template-parity.test.ts` pins * `ai.oas.json` to the scaffolder template byte for byte.) * * | Client | Path it sends | Rewritten to | * | ----------------------------------------- | ----------------------------------------------------------------- | ------------------------------------------------ | * | `AzureOpenAI` (from `openai`) | `/openai/deployments/{d}/chat/completions?api-version=…` | `/v1/chat/completions` | * | same | `/openai/deployments/{d}/embeddings` | `/v1/embeddings` | * | same | `/openai/responses[/{id}[/input_items]]` | `/v1/responses[/…]` | * | `@ai-sdk/azure`, `useDeploymentBasedUrls` | `/v1/deployments/{d}/{chat/completions,embeddings,responses}` | `/v1/{operation}` | * | `@ai-sdk/google-vertex/anthropic` | `/{model}:rawPredict` or `:streamRawPredict` | `/v1/messages`, `model` restored to the body | * | `@anthropic-ai/vertex-sdk` | `/projects/…/publishers/anthropic/models/{model}:rawPredict` | `/v1/messages`, `model` restored to the body | * | same | `…/publishers/anthropic/models/count-tokens:rawPredict` | `/v1/messages/count_tokens` | * * The deployment `{d}` is a gateway model reference, so it may contain one * slash (`azureai/my-gpt`) and never more; a Vertex `{model}` has two * (`vertexai/anthropic/claude-…`). The rules are bounded to those shapes, so a * path with extra segments in front of an operation is not recognized and * keeps its 404. The deployment is dropped from the path because the body * already carries `model`. The Vertex clients strip `model` from the body and * put it in the URL, so the shim moves it back; `anthropic_version`, which * they add, is accepted by the gateway. * * Credential: an `api-key` header — what the Azure clients send — becomes * `Authorization: Bearer` when no Authorization header is present, which is * the header `ai-gateway-auth-inbound` reads. Query strings such as * `api-version` are left alone; the gateway ignores them. * * Anything the shim does not recognize passes through with the path * unchanged, so canonical `/v1/...` and Bedrock `/model/...` requests keep * working on the same route, and unsupported operations still get the * gateway's own 404. * * This module is printed VERBATIM in the docs cookbook * `docs/ai-gateway/cookbooks/sdk-path-shim.mdx` (zuplo/docs). It is example * code a customer copies into their own gateway, not a runtime feature, and * `tests/sdk-shim.test.ts` is what keeps it working: change the two together. */ /** * Azure's two deployment-in-the-path grammars; the operation is captured. * * The deployment is bounded to one segment or `<assignment>/<name>` on * purpose: an unbounded `.+` would let * `/openai/deployments/d/audio/speech/chat/completions` masquerade as a chat * completion instead of keeping the 404 an unsupported operation deserves. */ const AZURE_DEPLOYMENT_PATHS = [ /^\/openai\/deployments\/[^/:]+(?:\/[^/:]+)?\/(chat\/completions|embeddings)$/, /^\/v1\/deployments\/[^/:]+(?:\/[^/:]+)?\/(chat\/completions|embeddings|responses)$/, ]; /** `AzureOpenAI` endpoints Azure serves without a deployment segment. */ const AZURE_BARE_PATH = /^\/openai\/(responses(?:\/.*)?)$/; /** * Vertex's Anthropic publisher surface, with or without the * `projects/{p}/locations/{l}/publishers/anthropic/models` prefix, and with an * optional `/v1` in front for a client whose `baseURL` already ends in `/v1`. * The model is one to three segments (`<assignment>/<publisher>/<model>`), * bounded for the same reason as the Azure deployment above. */ const VERTEX_ANTHROPIC_PATH = /^(?:\/v1)?(?:\/projects\/[^/]+\/locations\/[^/]+\/publishers\/anthropic\/models)?\/([^/:]+(?:\/[^/:]+){0,2}):(rawPredict|streamRawPredict)$/; /** * Client headers that describe the request body's bytes: its length, its * encoding, and any checksum. Dropped when the shim re-serializes the body, * because they would then describe bytes that no longer exist. The list matches * what the gateway itself strips before an upstream call * (`llm-translation/utils/forwarded-headers.ts`). */ const BODY_DESCRIBING_HEADERS = [ "content-length", "content-encoding", "content-md5", "content-digest", "digest", "repr-digest", ]; interface Rewrite { path: string; /** Set when the model was addressed in the URL and must go back in the body. */ modelFromPath?: string; } /** * Map the part of the path after the app id onto a gateway operation, or * return null when it is not one of the shimmed grammars. */ export function rewriteOperationPath(rest: string): Rewrite | null { for (const grammar of AZURE_DEPLOYMENT_PATHS) { const match = grammar.exec(rest); if (match) { return { path: `/v1/${match[1]}` }; } } const bare = AZURE_BARE_PATH.exec(rest); if (bare) { return { path: `/v1/${bare[1]}` }; } const vertex = VERTEX_ANTHROPIC_PATH.exec(rest); if (vertex) { let model: string; try { model = decodeURIComponent(vertex[1]); } catch { // A malformed percent escape is not a model reference. Leave the path // alone so the gateway answers its own 404 rather than a 500 from here. return null; } if (model === "count-tokens") { return { path: "/v1/messages/count_tokens" }; } return { path: "/v1/messages", modelFromPath: model }; } return null; } export default async function sdkPathShim( request: ZuploRequest, context: ZuploContext ): Promise<ZuploRequest> { const appId = request.params.app_id; if (typeof appId !== "string" || appId === "") { return request; } const url = new URL(request.url); const segments = url.pathname.split("/"); const appIndex = segments.indexOf(appId); if (appIndex === -1) { return request; } const rest = `/${segments.slice(appIndex + 1).join("/")}`; const rewrite = rewriteOperationPath(rest); // The loader requires the app id as the FIRST segment, so the rewritten path // is always `/{app_id}/…` whatever prefix this route was mounted on. const path = `/${appId}${rewrite?.path ?? rest}`; const headers = new Headers(request.headers); const apiKey = headers.get("api-key"); const movesKey = apiKey !== null && !headers.has("authorization"); if (movesKey) { headers.set("authorization", `Bearer ${apiKey}`); headers.delete("api-key"); } // Nothing to do for a canonical request on the main route: hand it on // untouched, body unread, so the shim costs it nothing. if (path === url.pathname && !movesKey) { return request; } // Only a Vertex rewrite needs the body — `model` was moved into the URL and // has to go back. Every other request keeps the body stream it arrived // with; nothing is buffered. let body: BodyInit | null = request.body; if ( rewrite?.modelFromPath !== undefined && request.method !== "GET" && request.method !== "HEAD" ) { // Bytes rather than text, so a body the shim ends up not changing (one // that already names its model, or is not JSON) goes on exactly as it // arrived, and the client's headers keep describing it. const bytes = await request.arrayBuffer(); body = bytes; let parsed: unknown; try { parsed = JSON.parse(new TextDecoder().decode(bytes)); } catch { parsed = undefined; } if (typeof parsed === "object" && parsed !== null) { const record = parsed as Record<string, unknown>; if (record.model === undefined) { record.model = rewrite.modelFromPath; body = JSON.stringify(record); // These are new bytes, so drop every header that described the old // ones — not just the length. for (const name of BODY_DESCRIBING_HEADERS) { headers.delete(name); } } } } if (path !== url.pathname) { context.log.debug( { from: url.pathname, to: path, policy: "sdk-path-shim-inbound" }, "SDK path shim rewrote the request path" ); } url.pathname = path; return new ZuploRequest(url.toString(), { method: request.method, headers, body, params: request.params, }); }

Install it in your gateway

The shim has to run before the gateway's configuration loader, so it goes on the catch-all route, not in an app's policy chain. The loader answers a 404 to any path whose segment after the app id isn't v1—before it loads the app—so a policy in an app's stored chain never sees a /openai/deployments/... or :rawPredict request. Mounting the shim first on the route is what gives it the chance to rewrite the path into a /v1/... shape the loader accepts.

Three steps, all in your gateway project:

  1. Copy the module

    Save the policy above as modules/sdk-path-shim-inbound.ts.

  2. Register the policy

    Add this entry to the policies array in config/policies.json:

    JSONCode
    { "name": "sdk-path-shim-inbound", "policyType": "custom-code-inbound", "handler": { "export": "default", "module": "$import(./modules/sdk-path-shim-inbound)" } }
  3. Put it first on the catch-all route

    In config/ai.oas.json, list sdk-path-shim-inbound first in the inbound policies of the AI Gateway catch-all route /:app_id/(.*), ahead of ai-gateway-configuration-loader-inbound:

    JSONCode
    "policies": { "inbound": [ "sdk-path-shim-inbound", "ai-gateway-configuration-loader-inbound", "ai-gateway-configuration-executor-inbound" ] }

    Older gateways list these two policies as ai-gateway-configuration-loader-v2-inbound and ai-gateway-configuration-executor-v2-inbound. Keep the names your route already uses and add the shim ahead of them.

The order isn't optional: the shim must run before the loader, for the reason above.

What each client needs

With the shim mounted, each client reaches the gateway by changing only its endpoint and credential. What each one needs on its side:

The AzureOpenAI class (openai package)

Construct it with your app's URL as endpoint and the app key as apiKey— nothing else changes:

TypeScriptCode
import { AzureOpenAI } from "openai"; const client = new AzureOpenAI({ endpoint: "https://my-gateway-main-2e18f50.zuplo.app/config_fe0a04972d2848e0a94ae4b8bcd1497e", apiVersion: "2024-10-21", apiKey: process.env.ZUPLO_APP_API_KEY, });

Chat completions, embeddings, and Responses (the client's /openai/responses path) all work. The shim strips the /openai/deployments/{deployment} and bare /openai prefixes, and moves the api-key header the class sends onto Authorization: Bearer.

@ai-sdk/azure with useDeploymentBasedUrls

The AI SDK Azure provider reaches the gateway without the shim as long as useDeploymentBasedUrls is off. Turn it on and the provider moves the deployment into the path; the shim rewrites it back:

TypeScriptCode
import { createAzure } from "@ai-sdk/azure"; const azure = createAzure({ baseURL: "https://my-gateway-main-2e18f50.zuplo.app/config_fe0a04972d2848e0a94ae4b8bcd1497e/v1", apiKey: process.env.ZUPLO_APP_API_KEY, useDeploymentBasedUrls: true, });

Here apiKey works—the shim maps the api-key header the provider sends onto Authorization: Bearer—and tokenProvider works too. azure.chat(id), the bare azure(id) (Responses), and azure.textEmbeddingModel(id) all reach the gateway.

@ai-sdk/google-vertex/anthropic

Point the provider's baseURL at your app and return the app key from generateAuthToken:

TypeScriptCode
import { createVertexAnthropic } from "@ai-sdk/google-vertex/anthropic"; const appKey = process.env.ZUPLO_APP_API_KEY; if (!appKey) { throw new Error("Set ZUPLO_APP_API_KEY to the app's API key"); } const vertexAnthropic = createVertexAnthropic({ baseURL: "https://my-gateway-main-2e18f50.zuplo.app/config_fe0a04972d2848e0a94ae4b8bcd1497e", generateAuthToken: async () => appKey, });

generateAuthToken is a public option, and it matters: without it the provider mints a Google access token, which is the wrong credential for the gateway. Its return type is Promise<string | null>, so read the key into a checked variable first rather than returning the environment lookup directly. The shim maps the provider's /{model}:rawPredict and :streamRawPredict paths onto /v1/messages and puts model back in the body. The anthropic_version field the provider adds is accepted by the gateway.

@anthropic-ai/vertex-sdk (AnthropicVertex)

This client resolves Google credentials before every call and has no option to skip that, so you override the credential rather than replace it. On a machine with Application Default Credentials, let it mint a Google token and override the header with the app key:

TypeScriptCode
import { AnthropicVertex } from "@anthropic-ai/vertex-sdk"; const client = new AnthropicVertex({ region: "global", projectId: "your-project", baseURL: "https://my-gateway-main-2e18f50.zuplo.app/config_fe0a04972d2848e0a94ae4b8bcd1497e", defaultHeaders: { Authorization: `Bearer ${process.env.ZUPLO_APP_API_KEY}`, }, });

Without ADC on the machine, pass an authClient whose getRequestHeaders() returns new Headers({ authorization: "Bearer <app key>" }) instead—a structural stand-in for a google-auth-library client, which is what the fixture test uses. Either way the shim sees a bearer app key.

region and projectId are required by the SDK but only end up in the URL, which the shim discards, so any values do. Messages, countTokens (the client's count-tokens:rawPredict path, mapped to /v1/messages/count_tokens), and streaming all work.

What keeps working

The shim is safe to mount on the main route because it rewrites only the grammars it recognizes:

  • Unrecognized paths are never rewritten. A canonical /v1/... request, a Bedrock /model/... request, and any operation the gateway doesn't serve keep their path—so your existing clients keep working, and an unsupported operation still gets the gateway's own 404 rather than a rewrite to something else. The credential move below is the one thing that applies regardless of path: a request carrying only api-key reaches the loader with that key moved to Authorization: Bearer, which is where the gateway reads it anyway.
  • It never replaces an existing Authorization header. The api-key→Authorization: Bearer move happens only when no Authorization header is present. Path and body rewrites still apply to such a request.
  • Query strings are left alone. Parameters such as api-version pass through untouched; the gateway ignores them.

What it doesn't cover

Two Gemini-native clients are out of reach: @ai-sdk/google-vertex's Gemini entry and @google/genai. They speak Gemini's native request and response format, which no URL rewrite can bridge—the gateway would have to translate the bodies, not just the path. Use an OpenAI client, or @ai-sdk/openai-compatible, for Gemini and Model Garden models, as the Vertex AI page explains.

Next steps

  • Custom Policies: the full custom-policy quickstart.
  • Using Azure AI: the supported Azure clients and how the gateway prices deployments.
  • Using Vertex AI: Claude on Vertex, and why the Gemini SDKs need an OpenAI client.
  • AI SDK: the Vercel AI SDK providers the gateway supports directly.
Last modified on