# Using Bedrock Runtime

**Bedrock Runtime** is
[Amazon Bedrock's native API](https://docs.aws.amazon.com/bedrock/latest/APIReference/API_Operations_Amazon_Bedrock_Runtime.html)—the
one the AWS SDKs call when you use `bedrock-runtime` from `@aws-sdk/*`, `boto3`,
or the Go and Java SDKs. Adding it as a provider lets those clients keep working
exactly as they are, with the gateway in the path for authentication, model
filtering, usage metering, and cost tracking.

This provider exists for one job: **you change the client's configuration, not
your code.** Point the client's `endpoint` at your app's URL, give it the app's
API key instead of AWS credentials, and every `ConverseCommand`,
`InvokeModelCommand`, and streaming call you already wrote keeps running.

Requests reach your own AWS account on your own AWS billing, signed by the
gateway with an IAM key pair you configure once.

## Bedrock Runtime or Bedrock Mantle?

Zuplo supports two Amazon Bedrock providers. Both reach your own AWS account—
what separates them is the API your client speaks.

|                   | Bedrock Runtime                                 | [Bedrock Mantle](./bedrock-mantle.mdx) |
| ----------------- | ----------------------------------------------- | -------------------------------------- |
| API your app uses | Bedrock's native API                            | OpenAI- or Anthropic-shaped            |
| Client            | an AWS SDK                                      | any OpenAI or Anthropic client         |
| Gateway path      | `/{app_id}/model/{modelId}/{operation}`         | `/{app_id}/v1/...`                     |
| Credential        | long-term IAM access key pair                   | long-term Bedrock API key (`ABSK`)     |
| Best for          | existing AWS SDK code you don't want to rewrite | new code, or one API across providers  |

Choose **Bedrock Runtime** when you already have AWS SDK code. Choose **Bedrock
Mantle** when you want Bedrock models through the same
[Universal API](./universal-api.mdx) as every other provider.

The two need separate provider configurations even for the same AWS account: a
provider holds exactly one credential, and a Bedrock API key can't be used as an
IAM access key pair.

## Supported operations

The gateway serves the four Bedrock Runtime operations at their AWS paths:

| Operation                       | Path                                           | Streaming |
| ------------------------------- | ---------------------------------------------- | --------- |
| `Converse`                      | `/model/{modelId}/converse`                    | No        |
| `ConverseStream`                | `/model/{modelId}/converse-stream`             | Yes       |
| `InvokeModel`                   | `/model/{modelId}/invoke`                      | No        |
| `InvokeModelWithResponseStream` | `/model/{modelId}/invoke-with-response-stream` | Yes       |

Your request and response bodies pass through unchanged, so anything these
operations support works—tool use, guardrail configuration,
`additionalModelRequestFields`, and model-native `InvokeModel` bodies the
gateway has no knowledge of. Streaming responses keep Bedrock's
`application/vnd.amazon.eventstream` framing, so the SDK's own decoder reads
them as it always has.

Your request headers reach Bedrock as well, including the `x-amzn-bedrock-*`
headers and the `x-amz-*` ones your SDK sets. The gateway removes only what it
owns: the `Authorization` header carrying your app's API key, the SigV4 headers
it replaces with its own signature (`x-amz-date`, `x-amz-content-sha256`, and
`x-amz-security-token`), cookies, the hop-by-hop headers, and the `zp-` and
`cf-` headers added on the way in. It sets `content-type` itself.

`CountTokens` and `InvokeModelWithBidirectionalStream` aren't served. Calling
either returns a `400 ValidationException` naming the operations that are.

:::note

A Bedrock Runtime provider serves only these native paths. It doesn't serve the
[Universal API](./universal-api.mdx) endpoints (`/v1/chat/completions`,
`/v1/messages`, `/v1/responses`, `/v1/embeddings`), and its models don't appear
in an app's `GET /v1/models` listing. Add [Bedrock Mantle](./bedrock-mantle.mdx)
alongside it if you want both.

:::

## Before you begin

You need:

- An AWS account with access to Amazon Bedrock in the region you plan to use.
- A **long-term IAM user access key pair**—an access key ID starting with `AKIA`
  and its secret. Temporary STS credentials (`ASIA…`, or any pair with a session
  token) are rejected: provider secrets are baked into the deployed gateway, so
  a credential that expires in hours would break the gateway between redeploys.
- Two IAM permissions on the models you plan to call, in that region:
  `bedrock:InvokeModel` and `bedrock:InvokeModelWithResponseStream`. That's the
  whole list—`bedrock:Converse` and `bedrock:ConverseStream` aren't IAM actions,
  and the Converse operations are authorized by the InvokeModel pair.
- An AI Gateway project in the Zuplo Portal, and an [app](./apps.mdx) to call
  from. The app page shows its API URL, and its API key lives on the app's **API
  Key** tab.

## Add the provider

Adding or editing providers requires the **Edit** permission, granted to Zuplo
account and project **Admins**—see
[Managing Providers](./managing-providers.mdx).

<Stepper>

1. Open
   [**Settings → AI Providers**](https://portal.zuplo.com/+/account/project/ai/settings/data-models)
   in your AI Gateway project in the Zuplo Portal.

1. Click the **Add Provider** button.

1. In the **AI Provider** list, select **Bedrock Runtime**.

1. Review the **Provider Name**, which fills in as `bedrockruntime`. You can
   replace it with your own name, but only now—the name is permanent after
   creation, and it's the prefix in every model reference.

1. In **AWS Region**, enter the lowercase code of the region you use Bedrock in,
   such as `us-east-1`. The gateway sends this provider's requests to
   `https://bedrock-runtime.<region>.amazonaws.com`—the standard endpoint
   family, not the FIPS or dual-stack ones. Regions in the China partition
   (`cn-`) aren't supported.

1. In **Access Key ID**, enter the key ID, which starts with `AKIA`. In **Secret
   Access Key**, enter the secret. AWS shows the secret only once, when the key
   is created.

1. Select the models to enable, then click **Create**.

</Stepper>

:::note

Saving provider settings triggers an automatic production deployment of your
gateway, because provider credentials are part of the deployed gateway. The
change is live once the deployment completes.

:::

## Point your AWS SDK at the gateway

Two configuration changes, and **both are required**:

1. Set `endpoint` to your app's API URL.
2. Make the client send the app's API key as a bearer token.

The second one matters more than it looks. An existing client usually has AWS
credentials available—in the environment, in a shared profile, or from an
instance role—and the SDK signs with SigV4 whenever it finds any. It would then
never send your gateway key, and the request would fail with a `401`. Setting
the auth scheme explicitly is what prevents that.

```ts
import {
  BedrockRuntimeClient,
  ConverseCommand,
} from "@aws-sdk/client-bedrock-runtime";

const apiKey = process.env.ZUPLO_APP_API_KEY;
if (!apiKey) {
  throw new Error("Set ZUPLO_APP_API_KEY to your app's API key.");
}

const client = new BedrockRuntimeClient({
  region: "us-east-1",
  // Your app's API URL—no /v1, no trailing path.
  endpoint:
    "https://my-gateway-main-2e18f50.zuplo.app/config_fe0a04972d2848e0a94ae4b8bcd1497e",
  // Send the app's API key as a bearer token instead of signing with SigV4.
  authSchemePreference: ["httpBearerAuth"],
  // `token` requires a string, so the check above is what makes this compile.
  token: { token: apiKey },
});

// Unchanged from here down.
const response = await client.send(
  new ConverseCommand({
    modelId: "bedrockruntime/us.anthropic.claude-haiku-4-5-20251001-v1:0",
    messages: [{ role: "user", content: [{ text: "Say hi" }] }],
  }),
);
```

Any AWS SDK that supports bearer-token authentication for Bedrock works the same
way—set the endpoint, supply the token, and make sure the client prefers bearer
auth over SigV4. Bearer-token support arrived in AWS SDK for JavaScript 3.840,
boto3 1.39, AWS SDK for Go 1.35, and AWS SDK for Java 2.31.74; use those
versions or later.

The gateway is the credential boundary: your client never holds AWS credentials,
and the gateway signs the upstream request with the IAM key pair you configured.

## Model references

Apps reference models as `providerName/modelId`, where `providerName` is the
name you gave the provider configuration:

```
bedrockruntime/us.anthropic.claude-haiku-4-5-20251001-v1:0
```

The gateway strips its own prefix and sends AWS the model ID exactly as you
wrote it. As with every other provider, the model must be one you enabled for
the provider. Cross-region inference profile IDs appear in the model list as
models of their own. A foundation-model or inference-profile ARN counts as the
model ID it contains, so it works when that model is enabled. The gateway
refuses any other model before it calls AWS, including a provisioned-throughput
ARN and an application inference profile, which don't name a model in the list.

:::tip{title="Many models need an inference profile ID"}

Bedrock refuses a bare model ID for models it serves only through cross-region
inference:

```
ValidationException: Invocation of model ID
anthropic.claude-haiku-4-5-20251001-v1:0 with on-demand throughput isn't
supported. Retry your request with the ID or ARN of an inference profile that
contains this model.
```

Use the inference profile ID—usually the model ID with a regional prefix such as
`us.`, `eu.`, or `apac.`. AWS lists them under
[Supported Regions and models for inference profiles](https://docs.aws.amazon.com/bedrock/latest/userguide/inference-profiles-support.html).

:::

## Pricing and budgets

The gateway prices Bedrock Runtime requests from the model catalog by exact
model ID, the same as every other provider. Cross-region inference profiles are
priced too: the `us.`, `eu.`, `apac.`, `au.`, `jp.`, `us-gov.`, and `global.`
IDs are catalog rows in their own right, each at its own rate. A
foundation-model or inference-profile ARN, such as
`arn:aws:bedrock:eu-central-1:123456789012:inference-profile/eu.anthropic.claude-opus-5-5`,
is priced as the ID it contains. Every model you can call has a catalog price,
so each request counts toward token, request, and spending
[budgets](./usage-limits.mdx).

When a budget is exhausted, these routes answer
`400 ServiceQuotaExceededException`—not the `429` the Universal API returns. AWS
SDKs treat any `429` as throttling and retry it, and a spent budget doesn't
refill between retries, so the gateway uses the status Bedrock itself uses for
an exhausted quota. Your client sees the refusal immediately instead of retrying
into it.

## Policies on Bedrock Runtime routes

These routes run your app's [policy chain](./policy-chains.mdx) like any other.
What changes is that the request body is Bedrock's shape rather than one the
gateway parses, so a policy that reads the body applies its `onUnknownShape`
setting instead. Each policy's default follows from its job: a guardrail fails
closed, while an observer or an optimization fails open.

| Policy                                                                                                                                  | On Bedrock Runtime routes                                                                                                                              |
| --------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------ |
| [Authentication](../policies/ai-gateway-auth-inbound.mdx)                                                                               | Works. Reads the app key from `Authorization: Bearer`                                                                                                  |
| [Model Filtering](../policies/ai-gateway-model-filtering-inbound.mdx)                                                                   | Works. Refuses a model outside the allow list before any AWS call                                                                                      |
| [Metering and budgets](../policies/ai-gateway-metering-inbound.mdx)                                                                     | Works. Meters tokens and cost on all four operations, streamed or not                                                                                  |
| [DLP](../policies/ai-gateway-dlp-inbound.mdx) and content guardrails                                                                    | **Blocks by default.** The gateway can't read these bodies, so the policy denies the request. Set `onUnknownShape` to `skip` to forward it uninspected |
| [Galileo](../policies/ai-gateway-galileo-tracing-inbound.mdx) and [Comet Opik](../policies/ai-gateway-opik-tracing-inbound.mdx) tracing | **Skipped by default.** The request goes through untraced. Set `onUnknownShape` to `deny` to make a trace a hard requirement                           |
| Semantic Cache                                                                                                                          | Skipped                                                                                                                                                |

The table lists every policy the gateway supports on these routes. It audits the
chain and refuses any other policy, such as a custom policy, Smart Router, or
[Fallback Model](./fallback.mdx), with a `400 ValidationException` that names
the policy. Remove it from the chain to use Bedrock Runtime through that app.

If you need content inspection on Bedrock traffic, use
[Bedrock's own guardrails](https://docs.aws.amazon.com/bedrock/latest/userguide/guardrails.html)
through `guardrailConfig` in the request body—the gateway forwards it untouched.

Errors the gateway raises on these routes carry an `x-amzn-ErrorType` header, so
your SDK surfaces them as typed exceptions rather than as deserialization
failures.

## Troubleshooting

**Every request returns `401`.** The client is signing with SigV4 instead of
sending the app's API key. Set `authSchemePreference` (or your SDK's equivalent)
to prefer bearer authentication, and confirm the token is the app's API key from
the app's **API Key** tab.

**`ValidationException` about on-demand throughput.** The model needs a
cross-region inference profile ID rather than the bare model ID—see
[Model references](#model-references).

**`403` from AWS that reads like a bad key.** Check that the IAM user has
`bedrock:InvokeModel` and `bedrock:InvokeModelWithResponseStream` on that model
_in the provider's region_, and that the model is enabled for your AWS account.

**The provider dialog rejects your key pair.** An access key ID starting with
`ASIA` is a temporary credential and is refused on purpose—it would expire while
baked into the deployed gateway. The dialog also rejects an access key ID that
isn't 16 to 128 uppercase letters and digits, a secret containing spaces or line
breaks, and either field filled in without the other.

**An error saying the model `is not included in model selections`.** The model
isn't enabled for the provider. Open the provider in
[**Settings → AI Providers**](https://portal.zuplo.com/+/account/project/ai/settings/data-models),
enable the model, and save. If the model is missing from the list, the gateway
can't serve it through this provider yet.

**A `400` naming the supported operations.** You called `CountTokens` or
`InvokeModelWithBidirectionalStream`, which the gateway doesn't serve.

**`ServiceQuotaExceededException` you can't find in AWS.** It's a Zuplo budget,
not an AWS service quota. An exhausted [budget](./usage-limits.mdx) refuses
these routes with that exception—check the app's usage limits before you open a
quota ticket with AWS.

**Your app's model list is empty of Bedrock models.** Expected—Bedrock Runtime
models don't appear in `GET /v1/models`, because that listing describes the
[Universal API](./universal-api.mdx), and this provider serves only Bedrock's
native paths.

## Next steps

- [AI Providers](./providers.mdx)—the capability matrix across every supported
  provider.
- [Using Bedrock Mantle](./bedrock-mantle.mdx)—the other Amazon Bedrock
  provider, for OpenAI- and Anthropic-shaped clients.
- [Managing Providers](./managing-providers.mdx)—edit models, keys, and the
  region, and understand when changes deploy.
- [AI Gateway Apps](./apps.mdx)—create the apps that call your models.
- [Model Filtering policy](../policies/ai-gateway-model-filtering-inbound.mdx)—control
  which models each app can call.
