# Cookbook: Point provider-native SDKs at the gateway

The AI Gateway serves every provider through the
[Universal API](../universal-api.mdx): OpenAI-shaped requests on
`/v1/chat/completions`, Anthropic-shaped ones on `/v1/messages`, and so on. Four
provider-native SDKs still refuse to talk to it—the `AzureOpenAI` class,
`@ai-sdk/azure` with `useDeploymentBasedUrls`,
`@ai-sdk/google-vertex/anthropic`, and `@anthropic-ai/vertex-sdk`—but not
because the gateway can't handle their requests. It already speaks each client's
body dialect. They fail only because each one hard-codes a provider-native URL
grammar—a deployment in the path, a `{model}:rawPredict` suffix—that the gateway
doesn't serve. A URL rewrite is all they need.

This recipe is a custom [inbound policy](../custom-policies.mdx) that does that
rewrite. Copy it into your own gateway and you own it: it's example code, not a
runtime feature, and these clients aren't among the supported clients documented
on the [Azure AI](../azure-ai.mdx) or [Vertex AI](../vertex-ai.mdx) pages. Model
references don't change—they stay the gateway's usual `providerName/deployment`
for Azure (`azureai/my-gpt`) and `providerName/publisher/model` for Vertex
(`vertexai/anthropic/claude-haiku-4-5@20251001`).

## The policy

Copy this module into your gateway as `modules/sdk-path-shim-inbound.ts`. It
rewrites the request path for the grammars it recognizes and leaves everything
else untouched; the header comment lists every path it maps, and
`tests/sdk-shim.test.ts` in zuplo/core is its executable half.

{/* prettier-ignore */}
```ts title="modules/sdk-path-shim-inbound.ts"
import type { ZuploContext } from "@zuplo/runtime";
import { ZuploRequest } from "@zuplo/runtime";

/**
 * SDK path shim: rewrites the URL grammars of four vendor SDKs onto the AI
 * Gateway's own endpoints, so a client that hard-codes a provider-native path
 * can be pointed at the gateway by changing its endpoint and credential only.
 *
 * Copy this module into your gateway, register it in `policies.json` as a
 * `custom-code-inbound` policy, and list it FIRST in the catch-all route's
 * inbound policies — before `ai-gateway-configuration-loader-inbound`. The
 * loader answers 404 to any path whose segment after the app id is not `v1`
 * before it loads the app, so a shim placed in an app's stored chain would
 * never run for these paths. (This fixture mounts it on a sibling route,
 * `/sdk-shim/:app_id/(.*)`, only because `tests/template-parity.test.ts` pins
 * `ai.oas.json` to the scaffolder template byte for byte.)
 *
 * | Client                                    | Path it sends                                                     | Rewritten to                                     |
 * | ----------------------------------------- | ----------------------------------------------------------------- | ------------------------------------------------ |
 * | `AzureOpenAI` (from `openai`)             | `/openai/deployments/{d}/chat/completions?api-version=…`          | `/v1/chat/completions`                           |
 * | same                                      | `/openai/deployments/{d}/embeddings`                              | `/v1/embeddings`                                 |
 * | same                                      | `/openai/responses[/{id}[/input_items]]`                          | `/v1/responses[/…]`                              |
 * | `@ai-sdk/azure`, `useDeploymentBasedUrls` | `/v1/deployments/{d}/{chat/completions,embeddings,responses}`     | `/v1/{operation}`                                |
 * | `@ai-sdk/google-vertex/anthropic`         | `/{model}:rawPredict` or `:streamRawPredict`                      | `/v1/messages`, `model` restored to the body     |
 * | `@anthropic-ai/vertex-sdk`                | `/projects/…/publishers/anthropic/models/{model}:rawPredict`      | `/v1/messages`, `model` restored to the body     |
 * | same                                      | `…/publishers/anthropic/models/count-tokens:rawPredict`           | `/v1/messages/count_tokens`                      |
 *
 * The deployment `{d}` is a gateway model reference, so it may contain one
 * slash (`azureai/my-gpt`) and never more; a Vertex `{model}` has two
 * (`vertexai/anthropic/claude-…`). The rules are bounded to those shapes, so a
 * path with extra segments in front of an operation is not recognized and
 * keeps its 404. The deployment is dropped from the path because the body
 * already carries `model`. The Vertex clients strip `model` from the body and
 * put it in the URL, so the shim moves it back; `anthropic_version`, which
 * they add, is accepted by the gateway.
 *
 * Credential: an `api-key` header — what the Azure clients send — becomes
 * `Authorization: Bearer` when no Authorization header is present, which is
 * the header `ai-gateway-auth-inbound` reads. Query strings such as
 * `api-version` are left alone; the gateway ignores them.
 *
 * Anything the shim does not recognize passes through with the path
 * unchanged, so canonical `/v1/...` and Bedrock `/model/...` requests keep
 * working on the same route, and unsupported operations still get the
 * gateway's own 404.
 *
 * This module is printed VERBATIM in the docs cookbook
 * `docs/ai-gateway/cookbooks/sdk-path-shim.mdx` (zuplo/docs). It is example
 * code a customer copies into their own gateway, not a runtime feature, and
 * `tests/sdk-shim.test.ts` is what keeps it working: change the two together.
 */

/**
 * Azure's two deployment-in-the-path grammars; the operation is captured.
 *
 * The deployment is bounded to one segment or `<assignment>/<name>` on
 * purpose: an unbounded `.+` would let
 * `/openai/deployments/d/audio/speech/chat/completions` masquerade as a chat
 * completion instead of keeping the 404 an unsupported operation deserves.
 */
const AZURE_DEPLOYMENT_PATHS = [
  /^\/openai\/deployments\/[^/:]+(?:\/[^/:]+)?\/(chat\/completions|embeddings)$/,
  /^\/v1\/deployments\/[^/:]+(?:\/[^/:]+)?\/(chat\/completions|embeddings|responses)$/,
];

/** `AzureOpenAI` endpoints Azure serves without a deployment segment. */
const AZURE_BARE_PATH = /^\/openai\/(responses(?:\/.*)?)$/;

/**
 * Vertex's Anthropic publisher surface, with or without the
 * `projects/{p}/locations/{l}/publishers/anthropic/models` prefix, and with an
 * optional `/v1` in front for a client whose `baseURL` already ends in `/v1`.
 * The model is one to three segments (`<assignment>/<publisher>/<model>`),
 * bounded for the same reason as the Azure deployment above.
 */
const VERTEX_ANTHROPIC_PATH =
  /^(?:\/v1)?(?:\/projects\/[^/]+\/locations\/[^/]+\/publishers\/anthropic\/models)?\/([^/:]+(?:\/[^/:]+){0,2}):(rawPredict|streamRawPredict)$/;

/**
 * Client headers that describe the request body's bytes: its length, its
 * encoding, and any checksum. Dropped when the shim re-serializes the body,
 * because they would then describe bytes that no longer exist. The list matches
 * what the gateway itself strips before an upstream call
 * (`llm-translation/utils/forwarded-headers.ts`).
 */
const BODY_DESCRIBING_HEADERS = [
  "content-length",
  "content-encoding",
  "content-md5",
  "content-digest",
  "digest",
  "repr-digest",
];

interface Rewrite {
  path: string;
  /** Set when the model was addressed in the URL and must go back in the body. */
  modelFromPath?: string;
}

/**
 * Map the part of the path after the app id onto a gateway operation, or
 * return null when it is not one of the shimmed grammars.
 */
export function rewriteOperationPath(rest: string): Rewrite | null {
  for (const grammar of AZURE_DEPLOYMENT_PATHS) {
    const match = grammar.exec(rest);
    if (match) {
      return { path: `/v1/${match[1]}` };
    }
  }

  const bare = AZURE_BARE_PATH.exec(rest);
  if (bare) {
    return { path: `/v1/${bare[1]}` };
  }

  const vertex = VERTEX_ANTHROPIC_PATH.exec(rest);
  if (vertex) {
    let model: string;
    try {
      model = decodeURIComponent(vertex[1]);
    } catch {
      // A malformed percent escape is not a model reference. Leave the path
      // alone so the gateway answers its own 404 rather than a 500 from here.
      return null;
    }
    if (model === "count-tokens") {
      return { path: "/v1/messages/count_tokens" };
    }
    return { path: "/v1/messages", modelFromPath: model };
  }

  return null;
}

export default async function sdkPathShim(
  request: ZuploRequest,
  context: ZuploContext
): Promise<ZuploRequest> {
  const appId = request.params.app_id;
  if (typeof appId !== "string" || appId === "") {
    return request;
  }

  const url = new URL(request.url);
  const segments = url.pathname.split("/");
  const appIndex = segments.indexOf(appId);
  if (appIndex === -1) {
    return request;
  }
  const rest = `/${segments.slice(appIndex + 1).join("/")}`;
  const rewrite = rewriteOperationPath(rest);

  // The loader requires the app id as the FIRST segment, so the rewritten path
  // is always `/{app_id}/…` whatever prefix this route was mounted on.
  const path = `/${appId}${rewrite?.path ?? rest}`;

  const headers = new Headers(request.headers);
  const apiKey = headers.get("api-key");
  const movesKey = apiKey !== null && !headers.has("authorization");
  if (movesKey) {
    headers.set("authorization", `Bearer ${apiKey}`);
    headers.delete("api-key");
  }

  // Nothing to do for a canonical request on the main route: hand it on
  // untouched, body unread, so the shim costs it nothing.
  if (path === url.pathname && !movesKey) {
    return request;
  }

  // Only a Vertex rewrite needs the body — `model` was moved into the URL and
  // has to go back. Every other request keeps the body stream it arrived
  // with; nothing is buffered.
  let body: BodyInit | null = request.body;
  if (
    rewrite?.modelFromPath !== undefined &&
    request.method !== "GET" &&
    request.method !== "HEAD"
  ) {
    // Bytes rather than text, so a body the shim ends up not changing (one
    // that already names its model, or is not JSON) goes on exactly as it
    // arrived, and the client's headers keep describing it.
    const bytes = await request.arrayBuffer();
    body = bytes;
    let parsed: unknown;
    try {
      parsed = JSON.parse(new TextDecoder().decode(bytes));
    } catch {
      parsed = undefined;
    }
    if (typeof parsed === "object" && parsed !== null) {
      const record = parsed as Record<string, unknown>;
      if (record.model === undefined) {
        record.model = rewrite.modelFromPath;
        body = JSON.stringify(record);
        // These are new bytes, so drop every header that described the old
        // ones — not just the length.
        for (const name of BODY_DESCRIBING_HEADERS) {
          headers.delete(name);
        }
      }
    }
  }

  if (path !== url.pathname) {
    context.log.debug(
      { from: url.pathname, to: path, policy: "sdk-path-shim-inbound" },
      "SDK path shim rewrote the request path"
    );
  }
  url.pathname = path;

  return new ZuploRequest(url.toString(), {
    method: request.method,
    headers,
    body,
    params: request.params,
  });
}
```

## Install it in your gateway

The shim has to run **before** the gateway's configuration loader, so it goes on
the catch-all route, not in an app's policy chain. The loader answers a `404` to
any path whose segment after the app id isn't `v1`—before it loads the app—so a
policy in an app's stored chain never sees a `/openai/deployments/...` or
`:rawPredict` request. Mounting the shim first on the route is what gives it the
chance to rewrite the path into a `/v1/...` shape the loader accepts.

Three steps, all in your gateway project:

<Stepper>

1. **Copy the module**

   Save the policy above as `modules/sdk-path-shim-inbound.ts`.

2. **Register the policy**

   Add this entry to the `policies` array in `config/policies.json`:

   ```json title="config/policies.json (one entry in the policies array)"
   {
     "name": "sdk-path-shim-inbound",
     "policyType": "custom-code-inbound",
     "handler": {
       "export": "default",
       "module": "$import(./modules/sdk-path-shim-inbound)"
     }
   }
   ```

3. **Put it first on the catch-all route**

   In `config/ai.oas.json`, list `sdk-path-shim-inbound` **first** in the
   inbound policies of the AI Gateway catch-all route `/:app_id/(.*)`, ahead of
   `ai-gateway-configuration-loader-inbound`:

   ```json title="config/ai.oas.json (the catch-all route's policies)"
   "policies": {
     "inbound": [
       "sdk-path-shim-inbound",
       "ai-gateway-configuration-loader-inbound",
       "ai-gateway-configuration-executor-inbound"
     ]
   }
   ```

   Older gateways list these two policies as
   `ai-gateway-configuration-loader-v2-inbound` and
   `ai-gateway-configuration-executor-v2-inbound`. Keep the names your route
   already uses and add the shim ahead of them.

</Stepper>

The order isn't optional: the shim must run before the loader, for the reason
above.

## What each client needs

With the shim mounted, each client reaches the gateway by changing only its
endpoint and credential. What each one needs on its side:

### The `AzureOpenAI` class (`openai` package)

Construct it with your app's URL as `endpoint` and the app key as `apiKey`—
nothing else changes:

```ts
import { AzureOpenAI } from "openai";

const client = new AzureOpenAI({
  endpoint:
    "https://my-gateway-main-2e18f50.zuplo.app/config_fe0a04972d2848e0a94ae4b8bcd1497e",
  apiVersion: "2024-10-21",
  apiKey: process.env.ZUPLO_APP_API_KEY,
});
```

Chat completions, embeddings, and Responses (the client's `/openai/responses`
path) all work. The shim strips the `/openai/deployments/{deployment}` and bare
`/openai` prefixes, and moves the `api-key` header the class sends onto
`Authorization: Bearer`.

### `@ai-sdk/azure` with `useDeploymentBasedUrls`

The [AI SDK Azure provider](../integrations/ai-sdk.mdx#azure-ai) reaches the
gateway without the shim as long as `useDeploymentBasedUrls` is off. Turn it on
and the provider moves the deployment into the path; the shim rewrites it back:

```ts
import { createAzure } from "@ai-sdk/azure";

const azure = createAzure({
  baseURL:
    "https://my-gateway-main-2e18f50.zuplo.app/config_fe0a04972d2848e0a94ae4b8bcd1497e/v1",
  apiKey: process.env.ZUPLO_APP_API_KEY,
  useDeploymentBasedUrls: true,
});
```

Here `apiKey` works—the shim maps the `api-key` header the provider sends onto
`Authorization: Bearer`—and `tokenProvider` works too. `azure.chat(id)`, the
bare `azure(id)` (Responses), and `azure.textEmbeddingModel(id)` all reach the
gateway.

### `@ai-sdk/google-vertex/anthropic`

Point the provider's `baseURL` at your app and return the app key from
`generateAuthToken`:

```ts
import { createVertexAnthropic } from "@ai-sdk/google-vertex/anthropic";

const appKey = process.env.ZUPLO_APP_API_KEY;
if (!appKey) {
  throw new Error("Set ZUPLO_APP_API_KEY to the app's API key");
}

const vertexAnthropic = createVertexAnthropic({
  baseURL:
    "https://my-gateway-main-2e18f50.zuplo.app/config_fe0a04972d2848e0a94ae4b8bcd1497e",
  generateAuthToken: async () => appKey,
});
```

`generateAuthToken` is a public option, and it matters: without it the provider
mints a Google access token, which is the wrong credential for the gateway. Its
return type is `Promise<string | null>`, so read the key into a checked variable
first rather than returning the environment lookup directly. The shim maps the
provider's `/{model}:rawPredict` and `:streamRawPredict` paths onto
`/v1/messages` and puts `model` back in the body. The `anthropic_version` field
the provider adds is accepted by the gateway.

### `@anthropic-ai/vertex-sdk` (`AnthropicVertex`)

This client resolves Google credentials before every call and has no option to
skip that, so you override the credential rather than replace it. On a machine
with Application Default Credentials, let it mint a Google token and override
the header with the app key:

```ts
import { AnthropicVertex } from "@anthropic-ai/vertex-sdk";

const client = new AnthropicVertex({
  region: "global",
  projectId: "your-project",
  baseURL:
    "https://my-gateway-main-2e18f50.zuplo.app/config_fe0a04972d2848e0a94ae4b8bcd1497e",
  defaultHeaders: {
    Authorization: `Bearer ${process.env.ZUPLO_APP_API_KEY}`,
  },
});
```

Without ADC on the machine, pass an `authClient` whose `getRequestHeaders()`
returns `new Headers({ authorization: "Bearer <app key>" })` instead—a
structural stand-in for a google-auth-library client, which is what the fixture
test uses. Either way the shim sees a bearer app key.

`region` and `projectId` are required by the SDK but only end up in the URL,
which the shim discards, so any values do. Messages, `countTokens` (the client's
`count-tokens:rawPredict` path, mapped to `/v1/messages/count_tokens`), and
streaming all work.

## What keeps working

The shim is safe to mount on the main route because it rewrites only the
grammars it recognizes:

- **Unrecognized paths are never rewritten.** A canonical `/v1/...` request, a
  Bedrock `/model/...` request, and any operation the gateway doesn't serve keep
  their path—so your existing clients keep working, and an unsupported operation
  still gets the gateway's own `404` rather than a rewrite to something else.
  The credential move below is the one thing that applies regardless of path: a
  request carrying only `api-key` reaches the loader with that key moved to
  `Authorization: Bearer`, which is where the gateway reads it anyway.
- **It never replaces an existing `Authorization` header.** The
  `api-key`→`Authorization: Bearer` move happens only when no `Authorization`
  header is present. Path and body rewrites still apply to such a request.
- **Query strings are left alone.** Parameters such as `api-version` pass
  through untouched; the gateway ignores them.

## What it doesn't cover

Two Gemini-native clients are out of reach: `@ai-sdk/google-vertex`'s Gemini
entry and `@google/genai`. They speak Gemini's native request _and response_
format, which no URL rewrite can bridge—the gateway would have to translate the
bodies, not just the path. Use an OpenAI client, or `@ai-sdk/openai-compatible`,
for Gemini and Model Garden models, as the
[Vertex AI page](../vertex-ai.mdx#call-the-models) explains.

## Next steps

- [Custom Policies](../custom-policies.mdx): the full custom-policy quickstart.
- [Using Azure AI](../azure-ai.mdx): the supported Azure clients and how the
  gateway prices deployments.
- [Using Vertex AI](../vertex-ai.mdx): Claude on Vertex, and why the Gemini SDKs
  need an OpenAI client.
- [AI SDK](../integrations/ai-sdk.mdx): the Vercel AI SDK providers the gateway
  supports directly.
