# Cache at the gateway

The Caching policy (`caching-inbound`) stores a response inside the gateway and
replays it for the next identical request. On a hit the policy answers from the
inbound pipeline and the backend never sees the request at all, which is the
whole point when the backend is the slow, rate-limited, or expensive part of the
system.

## When the gateway cache is the right layer

Gateway caching earns its place in three situations.

**No CDN sits in front of the gateway.** Traffic goes straight to Zuplo, so the
gateway is the only place a response can be reused. This is the common case for
internal APIs, partner APIs, and anything served directly from a Zuplo domain.

**The backend is slow, rate-limited, or billed per call.** A catalog search
endpoint that takes 900 ms of database time, or an upstream provider that allows
10 requests per second and charges per lookup, is worth protecting even when a
CDN also caches. A 60-second TTL on an endpoint receiving 200 requests per
minute turns 12,000 backend calls per hour into 60.

**The response must not leave the gateway boundary.** Data that policy or
compliance keeps off a third-party edge network can still be cached, because the
entry lives inside your Zuplo environment rather than in CDN storage.

:::tip

If a CDN already fronts your gateway, cache there first. A CDN hit is served
from a point of presence near the caller and never touches Zuplo at all, so it
is both cheaper and faster than a gateway hit. See
[Cache at the CDN](./cdn-caching.mdx) for the `cdn-cache-control-outbound`
policy. Use the gateway cache for what the CDN can't cover, not as a duplicate
layer.

:::

## How it works

The policy builds a cache key from the request, looks it up, and either answers
immediately or lets the request continue to the handler and stores the result on
the way out. By default the key covers the request method, the URL path, and the
query parameters, plus the `Authorization` header, any headers you name, and the
`cacheId` value. Two requests share an entry only when every one of those parts
matches.

<Diagram type="sequence" height="h-[420px]">
  <DiagramActor id="client" variant="yellow">
    Client
  </DiagramActor>
  <DiagramActor id="policy" variant="zuplo">
    Caching policy
  </DiagramActor>
  <DiagramActor id="backend" variant="green">
    Backend
  </DiagramActor>
  <DiagramMessage from="client" to="policy">
    GET /products?page=1 (miss)
  </DiagramMessage>
  <DiagramMessage from="policy" to="backend">
    Forward request
  </DiagramMessage>
  <DiagramMessage from="backend" to="policy">
    200 OK
  </DiagramMessage>
  <DiagramMessage from="policy" to="client">
    200 OK, response stored
  </DiagramMessage>
  <DiagramMessage from="client" to="policy">
    GET /products?page=1 (hit)
  </DiagramMessage>
  <DiagramMessage from="policy" to="client">
    200 OK from cache
  </DiagramMessage>
</Diagram>

The second request never reaches the backend. Neither does anything else the
route would normally do after the caching policy: a hit short-circuits the rest
of the pipeline.

## Add the policy to a route

Add the policy to `config/policies.json`. This example caches `GET` responses
for five minutes and keys them by locale.

```json title="config/policies.json"
{
  "policies": [
    {
      "name": "product-cache",
      "policyType": "caching-inbound",
      "handler": {
        "export": "CachingInboundPolicy",
        "module": "$import(@zuplo/runtime)",
        "options": {
          "expirationSecondsTtl": 300,
          "cacheHttpMethods": ["GET"],
          "statusCodes": [200, 404],
          "headers": ["Accept-Language"],
          "cacheId": "$env(CACHE_ID)",
          "dangerouslyIgnoreAuthorizationHeader": false
        }
      }
    }
  ]
}
```

Every option is optional. The defaults are:

| Option                                 | Default                               | What it controls                                 |
| -------------------------------------- | ------------------------------------- | ------------------------------------------------ |
| `expirationSecondsTtl`                 | `60`                                  | How long an entry stays valid, in seconds        |
| `cacheHttpMethods`                     | `["GET"]`                             | Which methods are cached                         |
| `statusCodes`                          | `[200, 206, 301, 302, 303, 404, 410]` | Which response status codes are stored           |
| `headers`                              | `[]`                                  | Extra request headers folded into the cache key  |
| `cacheId`                              | none                                  | An arbitrary string in the key, used for busting |
| `dangerouslyIgnoreAuthorizationHeader` | `false`                               | Whether to drop `Authorization` from the key     |

`cacheHttpMethods` accepts `GET`, `POST`, `PUT`, `PATCH`, `DELETE`, and `HEAD`.
Caching a method that changes state is almost always a mistake; leave the
default unless you have a read-only `POST` endpoint such as a search query.

Then reference the policy by name on the route:

```json title="config/routes.oas.json"
{
  "paths": {
    "/products": {
      "get": {
        "operationId": "list-products",
        "x-zuplo-route": {
          "handler": {
            "export": "urlForwardHandler",
            "module": "$import(@zuplo/runtime)",
            "options": { "baseUrl": "https://api.example.com" }
          },
          "policies": { "inbound": ["product-cache"] }
        }
      }
    }
  }
}
```

The string in `policies.inbound` must match the policy's `name` exactly. Put the
caching policy after authentication policies so that unauthenticated requests
are rejected before they can read from or write to the cache.

## Shape the cache key with headers

The `headers` option adds request headers to the key. Use it when the backend
returns a genuinely different body for different values of that header.

**`Accept`** — an endpoint that serves both `application/json` and `text/csv`
needs `"headers": ["Accept"]`. Without it, the first caller to ask for CSV
poisons the entry for every JSON caller for the rest of the TTL. Two media types
double the number of entries, which costs nothing.

**`Accept-Language`** — a catalog that localizes product names needs
`"headers": ["Accept-Language"]`. With five supported locales and 500 distinct
product URLs the cache holds 2,500 entries instead of 500, and the hit rate
within each locale is unchanged, because callers sharing a locale share an
entry.

Every header you add multiplies the key space by the number of distinct values
it carries in real traffic, which is why `User-Agent` is the classic mistake:
50,000 daily callers present tens of thousands of distinct strings, so nearly
every entry is written once and read never. The hit rate collapses to near zero
and the cache becomes pure overhead. To vary by device, normalize the variation
into a handful of values in an earlier policy and key on that instead.

## Cache authenticated responses

`dangerouslyIgnoreAuthorizationHeader` defaults to `false`, which means the
`Authorization` header is part of the cache key. Each bearer token or API key
gets its own entry, so one caller can never be served another caller's response.
Keep this default for anything user-specific.

Setting it to `true` removes `Authorization` from the key, and every authorized
caller shares one entry. That is legitimate in exactly one situation: the
response body is byte-identical for every caller who is allowed through — the
credential gates access but does not change the data. A shared reference table,
a public price list behind an API key, and a service status document all
qualify, and all benefit, because a single entry now absorbs traffic from every
consumer instead of one entry per token.

:::danger

If the response varies by caller in any way — a tenant ID in the payload, a
filtered result set, a personalized field — setting
`dangerouslyIgnoreAuthorizationHeader` to `true` serves the first caller's data
to everyone else on that route until the TTL expires. Nothing errors and nothing
logs; the disclosure is silent. Verify that two different authorized callers
receive identical bytes before you enable it.

:::

## Invalidate with cacheId

`cacheId` is an arbitrary string folded into the cache key. Change the string
and every previously stored key becomes unreachable at once. Drive it from an
environment variable so you can change it without editing code:

```json title="config/policies.json"
{
  "options": {
    "expirationSecondsTtl": 300,
    "cacheId": "$env(CACHE_ID)"
  }
}
```

Set `CACHE_ID` to a timestamp such as `2026-08-03-14-05` in your
[environment variables](../articles/environment-variables.mdx). To flush, update
the variable to a new value and redeploy. Every key the policy computes now
carries the new string, so no request matches an old entry. The orphaned entries
stay in storage until their TTL expires, but nothing serves them. A deployment
on its own flushes nothing — only a change to the value does. Give each policy
that needs independent invalidation its own variable; policies that share a
variable flush together.

:::warning

This is the entire invalidation story for the gateway cache. No purge API
exists, and no way to evict a single key. Plan around it: pick a TTL you are
willing to serve stale data for, and treat `cacheId` as the emergency lever
rather than part of a normal write path.

:::

Clients can also bust their own reads by adding a version or timestamp query
parameter. Query parameters are part of the key by default, so a new value
misses and fetches fresh data.

## Choose a TTL

`expirationSecondsTtl` defaults to 60 seconds. Pick a value from how fast the
underlying data actually changes, not from how fast you would like responses to
be — the TTL is the maximum staleness a caller can observe.

| Data volatility     | TTL           | Example endpoints                                                       |
| ------------------- | ------------- | ----------------------------------------------------------------------- |
| Static              | 3600 s and up | `/v1/countries`, `/v1/currencies`, `/v1/plans`, published documentation |
| Semi-dynamic        | 300–3600 s    | `/v1/products/{id}`, `/v1/venues`, `/v1/categories`, marketing content  |
| Frequently changing | 60–300 s      | `/v1/search?q=`, `/v1/showtimes`, `/v1/inventory`, dashboards           |

Anything that must reflect a write immediately — carts, seat holds, balances,
order status — does not belong in a response cache at any TTL. Start at the
short end of the range and lengthen it once you have measured the hit rate. A
TTL far below the interval between requests to a given key produces almost no
hits while still paying the write cost on every response.

## Limitations

**Cache entries are local to a zone.** Each Zuplo data center keeps its own
copy, so the first request for a key in a given region is always a miss even
when the same key is warm elsewhere. Expect a lower hit rate on globally
distributed traffic than a single-region test suggests.

**There is no selective purge.** Entries expire on their TTL or become
unreachable when `cacheId` changes. No API evicts a single key.

**Streaming responses are not cached**, and very large bodies are a poor fit —
the gateway has to hold the whole response to store it.

**A hit skips the rest of the pipeline.** Outbound policies do not run when the
caching policy answers from storage, so an outbound policy cannot be the sole
author of a header that must appear on every response.

:::caution{title="Do not combine with cdn-cache-control-outbound"}

Never put `caching-inbound` and
[`cdn-cache-control-outbound`](../policies/cdn-cache-control-outbound.mdx) on
the same route. `caching-inbound` answers a hit from the inbound pipeline, which
short-circuits the route before outbound policies run, so the CDN policy never
executes on a hit. The CDN then receives the copy that `caching-inbound` stored,
and `caching-inbound` sanitizes every copy it stores: it drops several CDN cache
headers (exactly which ones depends on the CDN) and overwrites `Cache-Control`
with its own `s-maxage` (its TTL, 60 seconds by default). The edge and client
TTLs collapse into one number. Misses run the full outbound stack and look
completely correct, so the defect appears only once the cache warms, and nothing
logs it. Pick one layer per route.

:::

## Specialized caches

Two policies key on payload shape rather than on bytes, and beat this policy on
the traffic they are built for:

- [GraphQL Cache](../policies/graphql-cache-inbound.mdx) normalizes a GraphQL
  document and its variables before hashing, so two semantically identical
  queries share an entry even when their bodies differ.
- [Semantic Cache](../policies/semantic-cache-inbound.mdx) matches requests by
  meaning rather than exact text, which is what makes caching viable in front of
  an LLM.

To cache a fragment of a response rather than all of it, or to drive the cache
from your own code, see
[Build a custom caching policy](./custom-caching-policy.mdx) and
[ZoneCache](../programmable-api/zone-cache.mdx).
