60-sec Short Edge & Caching 55 sec Failure mode + fix

CDN Cache Stampede

One cache expiry just took down your AI app.

Why it breaks

A cache only shields the origin while it's warm. When a hot key expires, every miss goes to the origin at once.

The fix

Coalesce misses into one origin fetch, serve stale-while-revalidate, add an origin shield and jitter TTLs.

Context

An AI product serves generated pages or answers: trending prompt results, AI summaries, model cards, "explore" feeds. Rendering one takes seconds of GPU time on the origin. A CDN caches the result at the edge with a short TTL so thousands of viewers share one render. Traffic is spiky: a viral page can jump to thousands of requests per second.

Architecture

ComponentRoleNotes
ViewersRequest the same popular URL10k req/s is illustrative
CDN edge (many PoPs)Caches the response with max-age=60Each PoP has its own cache
Origin (LLM / GPU service)Generates the page~2 s and one GPU slot per render (illustrative)
Origin capacitySized for cache misses, not total traffice.g. a few dozen concurrent renders

Request flow

  1. First request misses at the edge, the edge fetches from the origin, and the response is cached for 60 s.
  2. For the next 60 s, every request is a cache hit, so origin load is close to zero.
  3. At 60 s the object expires.

The failure mode

When the TTL runs out, every in-flight request for that key is a miss at the same moment, at every PoP. Without request collapsing, each miss becomes its own origin fetch, so thousands of identical 2-second GPU renders are requested at once. The origin's queue fills up, requests time out, the edge returns 5xx errors, and clients and load balancers retry, which adds more load. Because the cache never gets refilled, the stampede keeps going. This is also called dog-piling or the thundering herd.

Why it happens

The origin is sized for the cache's hit rate, but a synchronised expiry briefly takes the hit rate to zero for the hottest key.

The fix

  • Request coalescing / collapsed forwarding: the edge sends one origin request per key and makes the other misses wait for that response. Most CDNs support this, and some call it "request collapsing".
  • stale-while-revalidate / stale-if-error (RFC 5861): serve the expired copy right away while a single background request refreshes it. If the origin is failing, keep serving the stale copy.
  • Origin shield / tiered cache: all PoPs go through one mid-tier cache, so the origin sees one miss per key instead of one per PoP.
  • TTL jitter and early refresh: randomise expiry and refresh hot keys before they expire, so they don't all expire at once.
  • Protect the origin: set concurrency limits and load shedding, and have clients retry with backoff.

Trade-offs

  • Stale content: viewers may see a version that is up to the stale window out of date.
  • Coalescing adds latency for the waiting requests. If the single fetch fails, all of them see the failure unless stale-if-error is set.
  • An origin shield adds one hop, and the shield region becomes a dependency.

Numbers worth knowing

  • The traffic and render-time figures (10k req/s, 2 s per page, 60 s TTL) are illustrative.
  • stale-while-revalidate and stale-if-error are defined in RFC 5861.

Coming next in the series: Origin Shield explained

Found this useful?

Subscribe for the next episode, or share it with the person who owns this part of your stack.

Keep going