CDN Cache Stampede
One cache expiry just took down your AI app.
Why it breaks
A cache only shields the origin while it's warm. When a hot key expires, every miss goes to the origin at once.
The fix
Coalesce misses into one origin fetch, serve stale-while-revalidate, add an origin shield and jitter TTLs.
Context
An AI product serves generated pages or answers: trending prompt results, AI summaries, model cards, "explore" feeds. Rendering one takes seconds of GPU time on the origin. A CDN caches the result at the edge with a short TTL so thousands of viewers share one render. Traffic is spiky: a viral page can jump to thousands of requests per second.
Architecture
| Component | Role | Notes |
|---|---|---|
| Viewers | Request the same popular URL | 10k req/s is illustrative |
| CDN edge (many PoPs) | Caches the response with max-age=60 | Each PoP has its own cache |
| Origin (LLM / GPU service) | Generates the page | ~2 s and one GPU slot per render (illustrative) |
| Origin capacity | Sized for cache misses, not total traffic | e.g. a few dozen concurrent renders |
Request flow
- First request misses at the edge, the edge fetches from the origin, and the response is cached for 60 s.
- For the next 60 s, every request is a cache hit, so origin load is close to zero.
- At 60 s the object expires.
The failure mode
When the TTL runs out, every in-flight request for that key is a miss at the same moment, at every PoP. Without request collapsing, each miss becomes its own origin fetch, so thousands of identical 2-second GPU renders are requested at once. The origin's queue fills up, requests time out, the edge returns 5xx errors, and clients and load balancers retry, which adds more load. Because the cache never gets refilled, the stampede keeps going. This is also called dog-piling or the thundering herd.
Why it happens
The origin is sized for the cache's hit rate, but a synchronised expiry briefly takes the hit rate to zero for the hottest key.
The fix
- Request coalescing / collapsed forwarding: the edge sends one origin request per key and makes the other misses wait for that response. Most CDNs support this, and some call it "request collapsing".
stale-while-revalidate/stale-if-error(RFC 5861): serve the expired copy right away while a single background request refreshes it. If the origin is failing, keep serving the stale copy.- Origin shield / tiered cache: all PoPs go through one mid-tier cache, so the origin sees one miss per key instead of one per PoP.
- TTL jitter and early refresh: randomise expiry and refresh hot keys before they expire, so they don't all expire at once.
- Protect the origin: set concurrency limits and load shedding, and have clients retry with backoff.
Trade-offs
- Stale content: viewers may see a version that is up to the stale window out of date.
- Coalescing adds latency for the waiting requests. If the single fetch fails, all of them see the failure unless
stale-if-erroris set. - An origin shield adds one hop, and the shield region becomes a dependency.
Numbers worth knowing
- The traffic and render-time figures (10k req/s, 2 s per page, 60 s TTL) are illustrative.
stale-while-revalidateandstale-if-errorare defined in RFC 5861.
Coming next in the series: Origin Shield explained