Using Bedrock Runtime
Bedrock Runtime is
Amazon Bedrock's native API—the
one the AWS SDKs call when you use bedrock-runtime from @aws-sdk/*, boto3,
or the Go and Java SDKs. Adding it as a provider lets those clients keep working
exactly as they are, with the gateway in the path for authentication, model
filtering, usage metering, and cost tracking.
This provider exists for one job: you change the client's configuration, not
your code. Point the client's endpoint at your app's URL, give it the app's
API key instead of AWS credentials, and every ConverseCommand,
InvokeModelCommand, and streaming call you already wrote keeps running.
Requests reach your own AWS account on your own AWS billing, signed by the gateway with an IAM key pair you configure once.
Bedrock Runtime or Bedrock Mantle?
Zuplo supports two Amazon Bedrock providers. Both reach your own AWS account— what separates them is the API your client speaks.
| Bedrock Runtime | Bedrock Mantle | |
|---|---|---|
| API your app uses | Bedrock's native API | OpenAI- or Anthropic-shaped |
| Client | an AWS SDK | any OpenAI or Anthropic client |
| Gateway path | /{app_id}/model/{modelId}/{operation} | /{app_id}/v1/... |
| Credential | long-term IAM access key pair | long-term Bedrock API key (ABSK) |
| Best for | existing AWS SDK code you don't want to rewrite | new code, or one API across providers |
Choose Bedrock Runtime when you already have AWS SDK code. Choose Bedrock Mantle when you want Bedrock models through the same Universal API as every other provider.
The two need separate provider configurations even for the same AWS account: a provider holds exactly one credential, and a Bedrock API key can't be used as an IAM access key pair.
Supported operations
The gateway serves the four Bedrock Runtime operations at their AWS paths:
| Operation | Path | Streaming |
|---|---|---|
Converse | /model/{modelId}/converse | No |
ConverseStream | /model/{modelId}/converse-stream | Yes |
InvokeModel | /model/{modelId}/invoke | No |
InvokeModelWithResponseStream | /model/{modelId}/invoke-with-response-stream | Yes |
Your request and response bodies pass through unchanged, so anything these
operations support works—tool use, guardrail configuration,
additionalModelRequestFields, and model-native InvokeModel bodies the
gateway has no knowledge of. Streaming responses keep Bedrock's
application/vnd.amazon.eventstream framing, so the SDK's own decoder reads
them as it always has.
Your request headers reach Bedrock as well, including the x-amzn-bedrock-*
headers and the x-amz-* ones your SDK sets. The gateway removes only what it
owns: the Authorization header carrying your app's API key, the SigV4 headers
it replaces with its own signature (x-amz-date, x-amz-content-sha256, and
x-amz-security-token), cookies, the hop-by-hop headers, and the zp- and
cf- headers added on the way in. It sets content-type itself.
CountTokens and InvokeModelWithBidirectionalStream aren't served. Calling
either returns a 400 ValidationException naming the operations that are.
A Bedrock Runtime provider serves only these native paths. It doesn't serve the
Universal API endpoints (/v1/chat/completions,
/v1/messages, /v1/responses, /v1/embeddings), and its models don't appear
in an app's GET /v1/models listing. Add Bedrock Mantle
alongside it if you want both.
Before you begin
You need:
- An AWS account with access to Amazon Bedrock in the region you plan to use.
- A long-term IAM user access key pair—an access key ID starting with
AKIAand its secret. Temporary STS credentials (ASIA…, or any pair with a session token) are rejected: provider secrets are baked into the deployed gateway, so a credential that expires in hours would break the gateway between redeploys. - Two IAM permissions on the models you plan to call, in that region:
bedrock:InvokeModelandbedrock:InvokeModelWithResponseStream. That's the whole list—bedrock:Converseandbedrock:ConverseStreamaren't IAM actions, and the Converse operations are authorized by the InvokeModel pair. - An AI Gateway project in the Zuplo Portal, and an app to call from. The app page shows its API URL, and its API key lives on the app's API Key tab.
Add the provider
Adding or editing providers requires the Edit permission, granted to Zuplo account and project Admins—see Managing Providers.
-
Open Settings → AI Providers in your AI Gateway project in the Zuplo Portal.
-
In the AI Provider list, select Bedrock Runtime.
-
Review the Provider Name, which fills in as
bedrockruntime. You can replace it with your own name, but only now—the name is permanent after creation, and it's the prefix in every model reference. -
In AWS Region, enter the lowercase code of the region you use Bedrock in, such as
us-east-1. The gateway sends this provider's requests tohttps://bedrock-runtime.<region>.amazonaws.com—the standard endpoint family, not the FIPS or dual-stack ones. Regions in the China partition (cn-) aren't supported. -
In Access Key ID, enter the key ID, which starts with
AKIA. In Secret Access Key, enter the secret. AWS shows the secret only once, when the key is created. -
Select the models to enable, then click Create.
Saving provider settings triggers an automatic production deployment of your gateway, because provider credentials are part of the deployed gateway. The change is live once the deployment completes.
Point your AWS SDK at the gateway
Two configuration changes, and both are required:
- Set
endpointto your app's API URL. - Make the client send the app's API key as a bearer token.
The second one matters more than it looks. An existing client usually has AWS
credentials available—in the environment, in a shared profile, or from an
instance role—and the SDK signs with SigV4 whenever it finds any. It would then
never send your gateway key, and the request would fail with a 401. Setting
the auth scheme explicitly is what prevents that.
Code
Any AWS SDK that supports bearer-token authentication for Bedrock works the same way—set the endpoint, supply the token, and make sure the client prefers bearer auth over SigV4. Bearer-token support arrived in AWS SDK for JavaScript 3.840, boto3 1.39, AWS SDK for Go 1.35, and AWS SDK for Java 2.31.74; use those versions or later.
The gateway is the credential boundary: your client never holds AWS credentials, and the gateway signs the upstream request with the IAM key pair you configured.
Model references
Apps reference models as providerName/modelId, where providerName is the
name you gave the provider configuration:
Code
The gateway strips its own prefix and sends AWS the model ID exactly as you wrote it. As with every other provider, the model must be one you enabled for the provider. Cross-region inference profile IDs appear in the model list as models of their own. A foundation-model or inference-profile ARN counts as the model ID it contains, so it works when that model is enabled. The gateway refuses any other model before it calls AWS, including a provisioned-throughput ARN and an application inference profile, which don't name a model in the list.
Many models need an inference profile ID
Bedrock refuses a bare model ID for models it serves only through cross-region inference:
Code
Use the inference profile ID—usually the model ID with a regional prefix such as
us., eu., or apac.. AWS lists them under
Supported Regions and models for inference profiles.
Pricing and budgets
The gateway prices Bedrock Runtime requests from the model catalog by exact
model ID, the same as every other provider. Cross-region inference profiles are
priced too: the us., eu., apac., au., jp., us-gov., and global.
IDs are catalog rows in their own right, each at its own rate. A
foundation-model or inference-profile ARN, such as
arn:aws:bedrock:eu-central-1:123456789012:inference-profile/eu.anthropic.claude-opus-5-5,
is priced as the ID it contains. Every model you can call has a catalog price,
so each request counts toward token, request, and spending
budgets.
When a budget is exhausted, these routes answer
400 ServiceQuotaExceededException—not the 429 the Universal API returns. AWS
SDKs treat any 429 as throttling and retry it, and a spent budget doesn't
refill between retries, so the gateway uses the status Bedrock itself uses for
an exhausted quota. Your client sees the refusal immediately instead of retrying
into it.
Policies on Bedrock Runtime routes
These routes run your app's policy chain like any other.
What changes is that the request body is Bedrock's shape rather than one the
gateway parses, so a policy that reads the body applies its onUnknownShape
setting instead. Each policy's default follows from its job: a guardrail fails
closed, while an observer or an optimization fails open.
| Policy | On Bedrock Runtime routes |
|---|---|
| Authentication | Works. Reads the app key from Authorization: Bearer |
| Model Filtering | Works. Refuses a model outside the allow list before any AWS call |
| Metering and budgets | Works. Meters tokens and cost on all four operations, streamed or not |
| DLP and content guardrails | Blocks by default. The gateway can't read these bodies, so the policy denies the request. Set onUnknownShape to skip to forward it uninspected |
| Galileo and Comet Opik tracing | Skipped by default. The request goes through untraced. Set onUnknownShape to deny to make a trace a hard requirement |
| Semantic Cache | Skipped |
The table lists every policy the gateway supports on these routes. It audits the
chain and refuses any other policy, such as a custom policy, Smart Router, or
Fallback Model, with a 400 ValidationException that names
the policy. Remove it from the chain to use Bedrock Runtime through that app.
If you need content inspection on Bedrock traffic, use
Bedrock's own guardrails
through guardrailConfig in the request body—the gateway forwards it untouched.
Errors the gateway raises on these routes carry an x-amzn-ErrorType header, so
your SDK surfaces them as typed exceptions rather than as deserialization
failures.
Troubleshooting
Every request returns 401. The client is signing with SigV4 instead of
sending the app's API key. Set authSchemePreference (or your SDK's equivalent)
to prefer bearer authentication, and confirm the token is the app's API key from
the app's API Key tab.
ValidationException about on-demand throughput. The model needs a
cross-region inference profile ID rather than the bare model ID—see
Model references.
403 from AWS that reads like a bad key. Check that the IAM user has
bedrock:InvokeModel and bedrock:InvokeModelWithResponseStream on that model
in the provider's region, and that the model is enabled for your AWS account.
The provider dialog rejects your key pair. An access key ID starting with
ASIA is a temporary credential and is refused on purpose—it would expire while
baked into the deployed gateway. The dialog also rejects an access key ID that
isn't 16 to 128 uppercase letters and digits, a secret containing spaces or line
breaks, and either field filled in without the other.
An error saying the model is not included in model selections. The model
isn't enabled for the provider. Open the provider in
Settings → AI Providers,
enable the model, and save. If the model is missing from the list, the gateway
can't serve it through this provider yet.
A 400 naming the supported operations. You called CountTokens or
InvokeModelWithBidirectionalStream, which the gateway doesn't serve.
ServiceQuotaExceededException you can't find in AWS. It's a Zuplo budget,
not an AWS service quota. An exhausted budget refuses
these routes with that exception—check the app's usage limits before you open a
quota ticket with AWS.
Your app's model list is empty of Bedrock models. Expected—Bedrock Runtime
models don't appear in GET /v1/models, because that listing describes the
Universal API, and this provider serves only Bedrock's
native paths.
Next steps
- AI Providers—the capability matrix across every supported provider.
- Using Bedrock Mantle—the other Amazon Bedrock provider, for OpenAI- and Anthropic-shaped clients.
- Managing Providers—edit models, keys, and the region, and understand when changes deploy.
- AI Gateway Apps—create the apps that call your models.
- Model Filtering policy—control which models each app can call.