Providers

Using Bedrock Runtime

Bedrock Runtime is Amazon Bedrock's native API—the one the AWS SDKs call when you use bedrock-runtime from @aws-sdk/*, boto3, or the Go and Java SDKs. Adding it as a provider lets those clients keep working exactly as they are, with the gateway in the path for authentication, model filtering, usage metering, and cost tracking.

This provider exists for one job: you change the client's configuration, not your code. Point the client's endpoint at your app's URL, give it the app's API key instead of AWS credentials, and every ConverseCommand, InvokeModelCommand, and streaming call you already wrote keeps running.

Requests reach your own AWS account on your own AWS billing, signed by the gateway with an IAM key pair you configure once.

Bedrock Runtime or Bedrock Mantle?

Zuplo supports two Amazon Bedrock providers. Both reach your own AWS account— what separates them is the API your client speaks.

Bedrock RuntimeBedrock Mantle
API your app usesBedrock's native APIOpenAI- or Anthropic-shaped
Clientan AWS SDKany OpenAI or Anthropic client
Gateway path/{app_id}/model/{modelId}/{operation}/{app_id}/v1/...
Credentiallong-term IAM access key pairlong-term Bedrock API key (ABSK)
Best forexisting AWS SDK code you don't want to rewritenew code, or one API across providers

Choose Bedrock Runtime when you already have AWS SDK code. Choose Bedrock Mantle when you want Bedrock models through the same Universal API as every other provider.

The two need separate provider configurations even for the same AWS account: a provider holds exactly one credential, and a Bedrock API key can't be used as an IAM access key pair.

Supported operations

The gateway serves the four Bedrock Runtime operations at their AWS paths:

OperationPathStreaming
Converse/model/{modelId}/converseNo
ConverseStream/model/{modelId}/converse-streamYes
InvokeModel/model/{modelId}/invokeNo
InvokeModelWithResponseStream/model/{modelId}/invoke-with-response-streamYes

Your request and response bodies pass through unchanged, so anything these operations support works—tool use, guardrail configuration, additionalModelRequestFields, and model-native InvokeModel bodies the gateway has no knowledge of. Streaming responses keep Bedrock's application/vnd.amazon.eventstream framing, so the SDK's own decoder reads them as it always has.

Your request headers reach Bedrock as well, including the x-amzn-bedrock-* headers and the x-amz-* ones your SDK sets. The gateway removes only what it owns: the Authorization header carrying your app's API key, the SigV4 headers it replaces with its own signature (x-amz-date, x-amz-content-sha256, and x-amz-security-token), cookies, the hop-by-hop headers, and the zp- and cf- headers added on the way in. It sets content-type itself.

CountTokens and InvokeModelWithBidirectionalStream aren't served. Calling either returns a 400 ValidationException naming the operations that are.

A Bedrock Runtime provider serves only these native paths. It doesn't serve the Universal API endpoints (/v1/chat/completions, /v1/messages, /v1/responses, /v1/embeddings), and its models don't appear in an app's GET /v1/models listing. Add Bedrock Mantle alongside it if you want both.

Before you begin

You need:

  • An AWS account with access to Amazon Bedrock in the region you plan to use.
  • A long-term IAM user access key pair—an access key ID starting with AKIA and its secret. Temporary STS credentials (ASIA…, or any pair with a session token) are rejected: provider secrets are baked into the deployed gateway, so a credential that expires in hours would break the gateway between redeploys.
  • Two IAM permissions on the models you plan to call, in that region: bedrock:InvokeModel and bedrock:InvokeModelWithResponseStream. That's the whole list—bedrock:Converse and bedrock:ConverseStream aren't IAM actions, and the Converse operations are authorized by the InvokeModel pair.
  • An AI Gateway project in the Zuplo Portal, and an app to call from. The app page shows its API URL, and its API key lives on the app's API Key tab.

Add the provider

Adding or editing providers requires the Edit permission, granted to Zuplo account and project Admins—see Managing Providers.

  1. Open Settings → AI Providers in your AI Gateway project in the Zuplo Portal.

  2. Click the Add Provider button.

  3. In the AI Provider list, select Bedrock Runtime.

  4. Review the Provider Name, which fills in as bedrockruntime. You can replace it with your own name, but only now—the name is permanent after creation, and it's the prefix in every model reference.

  5. In AWS Region, enter the lowercase code of the region you use Bedrock in, such as us-east-1. The gateway sends this provider's requests to https://bedrock-runtime.<region>.amazonaws.com—the standard endpoint family, not the FIPS or dual-stack ones. Regions in the China partition (cn-) aren't supported.

  6. In Access Key ID, enter the key ID, which starts with AKIA. In Secret Access Key, enter the secret. AWS shows the secret only once, when the key is created.

  7. Select the models to enable, then click Create.

Saving provider settings triggers an automatic production deployment of your gateway, because provider credentials are part of the deployed gateway. The change is live once the deployment completes.

Point your AWS SDK at the gateway

Two configuration changes, and both are required:

  1. Set endpoint to your app's API URL.
  2. Make the client send the app's API key as a bearer token.

The second one matters more than it looks. An existing client usually has AWS credentials available—in the environment, in a shared profile, or from an instance role—and the SDK signs with SigV4 whenever it finds any. It would then never send your gateway key, and the request would fail with a 401. Setting the auth scheme explicitly is what prevents that.

TypeScriptCode
import { BedrockRuntimeClient, ConverseCommand, } from "@aws-sdk/client-bedrock-runtime"; const apiKey = process.env.ZUPLO_APP_API_KEY; if (!apiKey) { throw new Error("Set ZUPLO_APP_API_KEY to your app's API key."); } const client = new BedrockRuntimeClient({ region: "us-east-1", // Your app's API URL—no /v1, no trailing path. endpoint: "https://my-gateway-main-2e18f50.zuplo.app/config_fe0a04972d2848e0a94ae4b8bcd1497e", // Send the app's API key as a bearer token instead of signing with SigV4. authSchemePreference: ["httpBearerAuth"], // `token` requires a string, so the check above is what makes this compile. token: { token: apiKey }, }); // Unchanged from here down. const response = await client.send( new ConverseCommand({ modelId: "bedrockruntime/us.anthropic.claude-haiku-4-5-20251001-v1:0", messages: [{ role: "user", content: [{ text: "Say hi" }] }], }), );

Any AWS SDK that supports bearer-token authentication for Bedrock works the same way—set the endpoint, supply the token, and make sure the client prefers bearer auth over SigV4. Bearer-token support arrived in AWS SDK for JavaScript 3.840, boto3 1.39, AWS SDK for Go 1.35, and AWS SDK for Java 2.31.74; use those versions or later.

The gateway is the credential boundary: your client never holds AWS credentials, and the gateway signs the upstream request with the IAM key pair you configured.

Model references

Apps reference models as providerName/modelId, where providerName is the name you gave the provider configuration:

Code
bedrockruntime/us.anthropic.claude-haiku-4-5-20251001-v1:0

The gateway strips its own prefix and sends AWS the model ID exactly as you wrote it. As with every other provider, the model must be one you enabled for the provider. Cross-region inference profile IDs appear in the model list as models of their own. A foundation-model or inference-profile ARN counts as the model ID it contains, so it works when that model is enabled. The gateway refuses any other model before it calls AWS, including a provisioned-throughput ARN and an application inference profile, which don't name a model in the list.

Many models need an inference profile ID

Bedrock refuses a bare model ID for models it serves only through cross-region inference:

Code
ValidationException: Invocation of model ID anthropic.claude-haiku-4-5-20251001-v1:0 with on-demand throughput isn't supported. Retry your request with the ID or ARN of an inference profile that contains this model.

Use the inference profile ID—usually the model ID with a regional prefix such as us., eu., or apac.. AWS lists them under Supported Regions and models for inference profiles.

Pricing and budgets

The gateway prices Bedrock Runtime requests from the model catalog by exact model ID, the same as every other provider. Cross-region inference profiles are priced too: the us., eu., apac., au., jp., us-gov., and global. IDs are catalog rows in their own right, each at its own rate. A foundation-model or inference-profile ARN, such as arn:aws:bedrock:eu-central-1:123456789012:inference-profile/eu.anthropic.claude-opus-5-5, is priced as the ID it contains. Every model you can call has a catalog price, so each request counts toward token, request, and spending budgets.

When a budget is exhausted, these routes answer 400 ServiceQuotaExceededException—not the 429 the Universal API returns. AWS SDKs treat any 429 as throttling and retry it, and a spent budget doesn't refill between retries, so the gateway uses the status Bedrock itself uses for an exhausted quota. Your client sees the refusal immediately instead of retrying into it.

Policies on Bedrock Runtime routes

These routes run your app's policy chain like any other. What changes is that the request body is Bedrock's shape rather than one the gateway parses, so a policy that reads the body applies its onUnknownShape setting instead. Each policy's default follows from its job: a guardrail fails closed, while an observer or an optimization fails open.

PolicyOn Bedrock Runtime routes
AuthenticationWorks. Reads the app key from Authorization: Bearer
Model FilteringWorks. Refuses a model outside the allow list before any AWS call
Metering and budgetsWorks. Meters tokens and cost on all four operations, streamed or not
DLP and content guardrailsBlocks by default. The gateway can't read these bodies, so the policy denies the request. Set onUnknownShape to skip to forward it uninspected
Galileo and Comet Opik tracingSkipped by default. The request goes through untraced. Set onUnknownShape to deny to make a trace a hard requirement
Semantic CacheSkipped

The table lists every policy the gateway supports on these routes. It audits the chain and refuses any other policy, such as a custom policy, Smart Router, or Fallback Model, with a 400 ValidationException that names the policy. Remove it from the chain to use Bedrock Runtime through that app.

If you need content inspection on Bedrock traffic, use Bedrock's own guardrails through guardrailConfig in the request body—the gateway forwards it untouched.

Errors the gateway raises on these routes carry an x-amzn-ErrorType header, so your SDK surfaces them as typed exceptions rather than as deserialization failures.

Troubleshooting

Every request returns 401. The client is signing with SigV4 instead of sending the app's API key. Set authSchemePreference (or your SDK's equivalent) to prefer bearer authentication, and confirm the token is the app's API key from the app's API Key tab.

ValidationException about on-demand throughput. The model needs a cross-region inference profile ID rather than the bare model ID—see Model references.

403 from AWS that reads like a bad key. Check that the IAM user has bedrock:InvokeModel and bedrock:InvokeModelWithResponseStream on that model in the provider's region, and that the model is enabled for your AWS account.

The provider dialog rejects your key pair. An access key ID starting with ASIA is a temporary credential and is refused on purpose—it would expire while baked into the deployed gateway. The dialog also rejects an access key ID that isn't 16 to 128 uppercase letters and digits, a secret containing spaces or line breaks, and either field filled in without the other.

An error saying the model is not included in model selections. The model isn't enabled for the provider. Open the provider in Settings → AI Providers, enable the model, and save. If the model is missing from the list, the gateway can't serve it through this provider yet.

A 400 naming the supported operations. You called CountTokens or InvokeModelWithBidirectionalStream, which the gateway doesn't serve.

ServiceQuotaExceededException you can't find in AWS. It's a Zuplo budget, not an AWS service quota. An exhausted budget refuses these routes with that exception—check the app's usage limits before you open a quota ticket with AWS.

Your app's model list is empty of Bedrock models. Expected—Bedrock Runtime models don't appear in GET /v1/models, because that listing describes the Universal API, and this provider serves only Bedrock's native paths.

Next steps

Last modified on