60-sec Short Edge & Caching 56 sec Failure mode + fix

Edge Rendering Trap

Edge rendering made your AI app slower.

Video premiering soon on @AI.JoinDev Subscribe to get it first

Why it's slow

Edge compute only helps when the work is at the edge. Every await on distant data pays a full ocean round trip.

The fix

Render next to your data in one region. Run fetches in parallel and stream a cached shell from the edge.

Context

A Next.js AI chat app: a signed-in user opens a chat page that shows their history, retrieves context from a vector store and asks an LLM for a first suggestion. The team set export const runtime = 'edge' on the page so it would render "close to the user". The users are global. The database, the vector index and the LLM provider's endpoint all sit in one region, us-east-1, because that's where the data lives.

Architecture

ComponentRoleNotes
User (Sydney)Opens the chat pageAny user far from the data region
Edge render (Sydney PoP)Runs the page's server render next to the userruntime = 'edge', now deprecated in Next.js 16
Database + vector index (us-east-1)Session, chat history, retrievalSingle-region primary; not globally replicated
LLM API (us-east-1)Generates the first answerSame region as the data

Request flow

  1. The request reaches the nearest edge location, and the page's server component starts rendering there.
  2. It awaits the session check, then the chat history, then the retrieval query, then the LLM call, each one after the one before it.
  3. Only after the last await does the HTML start streaming back to the user.

The failure mode

Each await is a round trip from Sydney to Virginia and back, roughly 200 ms (illustrative) before the service even does any work. Four sequential awaits cost about 800 ms of pure network time, so the user sees a blank page for about a second. The same render in us-east-1 would pay only a few milliseconds per hop, plus one Sydney round trip for the HTML itself. So "rendering at the edge" made the page slower than rendering in one region. It's easy to miss, because it looks fast in a demo run from the same continent as the database, and because the edge function's own compute time is tiny.

Why it happens

Latency is dominated by distance × number of sequential round trips. Moving compute next to the user shortens one hop (user → server) but lengthens every server → data hop, and a render usually has several of those.

The fix

  • Render next to your data: use the Node.js runtime (the default; Next.js 16 deprecates runtime = 'edge' for pages and routes) and run the function in the database's region (on Vercel, the project's function region or the preferredRegion segment config).
  • Kill the waterfall: start independent fetches together (Promise.all), and move work that depends on nothing (auth, history) out of the LLM's critical path.
  • Stream a shell from the edge: keep the static layout cached at the CDN and stream the dynamic parts in with <Suspense>. With Partial Prerendering (Cache Components in Next.js 16) the user sees the shell instantly while the regional render fills in the rest.
  • Keep the edge for edge-shaped work: redirects, geolocation, A/B bucketing and cheap auth checks in Proxy (previously Middleware) don't need the database.

Trade-offs

  • Far-away users pay one long round trip for the dynamic HTML. That's a single hop, not four.
  • Replicating data to many regions (read replicas, edge KV) can make edge rendering pay off, but it brings replication lag and consistency work.
  • Streaming and Partial Prerendering need the page split into static and dynamic parts, and errors show up mid-stream.

Numbers worth knowing

  • The 200 ms Sydney ↔ us-east-1 round trip and the four sequential awaits are illustrative; measure your own region-to-region RTT before quoting a figure.
  • Next.js marks export const runtime = 'edge' as deprecated (the "Edge Runtime Deprecated" message in the Next.js docs); the Node.js runtime is the default.
  • In 2024 Vercel moved its own edge-rendered pages back to Node.js. Lee Robinson's explanation: most data isn't globally replicated, so compute belongs next to the database.

Coming next in the series: Partial Prerendering

Found this useful?

Subscribe for the next episode, or share it with the person who owns this part of your stack.

Keep going