AI Gateway overview
Zuplo's AI Gateway acts as an intelligent proxy layer that sits between your apps, your people, and LLM providers like OpenAI, Anthropic, Google, Mistral, and xAI. Instead of your apps communicating directly with these providers, all requests flow through the Zuplo AI Gateway, which streams responses while applying policies, controls, and monitoring.
Each piece of software that calls the gateway registers as an app—your support chatbot is one app; an internal coding agent is another. A single AI Gateway holds as many apps as you need, and you control exactly which policies run for each one. Every app has its own ordered policy chain, so you can give the customer-facing chatbot an app-specific budget and guardrails while leaving those controls out of the internal agent's chain. Beyond the built-in policies, you can write your own in TypeScript, and apps add them to their chains like any other policy. Chain changes apply within about a minute, with no redeploy. See Apps.
People call the gateway too. Instead of creating an app for each person, every AI Gateway project has a User App: people create their own personal API key and use it from tools such as Claude Code or Codex. Each person draws from their own budget, and admins see who spent what. See User Apps.
Key Benefits
Provider Independence: Switch between LLM providers (OpenAI, Anthropic, Google, Mistral, xAI, and more) dynamically without modifying app code. Configure model access through the gateway rather than hard coding it into your apps.
Cost Control: Set independent spending limits at the gateway, pool, and app levels. Configure hourly, daily, weekly, or monthly rules that warn or block when usage reaches a limit.
Security & Compliance: Apply guardrails to detect and block prompt injection attempts and prevent PII leakage in both requests and responses through integrated AI firewall policies.
Self-Service Access: Developers can create apps, and anyone on the project can create a personal API key, without needing direct access to provider API keys. Administrators configure providers once, and apps and people use them securely.
Extensibility: Write custom policies in TypeScript—content filters, dynamic model routing, plan-based limits—and let apps add them to their chains like any built-in policy.
Performance Optimization: Enable semantic caching to identify and return cached responses for similar prompts, reducing costs and improving response times.
Full Observability: Real-time dashboards show request counts, token usage, time-to-first-byte metrics, and spending patterns across your gateway.
How It Works
Your code sends its LLM requests to your AI Gateway instead of directly to a provider. Each app has its own URL, API key, and policy chain, so one gateway serves apps with completely different configurations. People send requests to the User App's URL with their personal key instead.
Each request arrives at the URL of the app it belongs to. The gateway identifies that app, runs its policy chain (model access, app-specific cost controls, security guardrails), routes to the selected LLM provider, and streams the response back. Throughout this process, the gateway captures metrics without exposing underlying provider credentials.
Every app-specific control is opt-in. An app whose policy chain is empty runs no app-selected controls: each request names its own model, and no app-specific model restrictions, budgets, guardrails, or caching apply. Gateway and pool usage limits still apply.
Core Features
Multi-Provider Support
Configure multiple LLM providers within a single AI Gateway project. Supported
providers include OpenAI, Anthropic, Google, Mistral, xAI, Microsoft Azure
(through Azure AI), Amazon Bedrock (through
Bedrock Mantle), Google Cloud
(Vertex AI), OpenRouter, TypeSafe's
System One model (Jev), and OpenAI-compatible custom providers. See
AI Providers for the full list of providers and supported
capabilities. Apps reference models as providerName/model—for example
openai/gpt-6-luna—so a single app can use models from several providers.
Source-Controlled Gateway
A new AI Gateway project deploys without a Git repository, so you can configure providers, pools, and apps right away. Connect a Git repository when you want to write custom policies or review gateway changes in Git. After you connect, the default branch is what production runs, and new template routes and policies reach the gateway only when you add them to the repository.
Per-App Policy Chains
Each app runs its own ordered chain of policies—model filtering, fallback models, app-specific budgets, semantic caching, guardrails, tracing, and any custom policies declared in the repository. Pool policy templates give new apps a consistent starting pipeline. See Policy Chains and Custom Policies.
Pool Hierarchy & Budgets
Organize apps into pools with hierarchical structures. Set a budget at each node. Each node totals its own usage and all descendant usage, and any exceeded Block limit rejects the request:
- Gateway: Limits across the Zuplo project (for example, $1,000/day).
- Pools: Limits across the pool, its sub-pools, and their apps (for example, $500/day for the Credit Pool).
- Apps: Limits for one app's traffic.
The User App has a per-person budget instead, such as $50 per day for each person. Gateway limits apply to it as well. See Usage Limits.
App Configuration
Each app gets its own:
- Unique Gateway URL: Single endpoint regardless of underlying provider
- API Key: Zuplo-managed key that never exposes provider credentials
- Policy Chain: Model access, app-specific budgets, caching, guardrails, and custom policies, applied in the order the app chooses
Zuplo accounts and projects
A Zuplo account is the top-level container for your members and projects. Each Zuplo project belongs to one account, and an AI Gateway is a type of Zuplo project. The providers, pools, and apps you configure for a gateway all belong to that project.
Account roles can grant access across the projects in a Zuplo account, while project roles grant access to one Zuplo project. AI Gateway pool roles add access to a pool and its apps within that project. The AI User project role grants only what someone needs to call the User App with their own key. See Role Permissions for the exact permissions each role grants.
Use Cases
- Multi-tenant AI Apps: Enforce spending limits per customer or pool
- Everyday AI for your people: Give everyone a personal API key with their own budget for coding assistants and chat tools
- Agent Development: Build AI agents that can switch providers without code changes
- Cost Management: Control and monitor LLM spending across your gateway
- Security Compliance: Ensure PII and prompt injection protection across all LLM interactions
- Performance: Reduce costs and latency with semantic caching for common queries