Cinematic deep dive More system design 7 min Failure mode + fix

Build Web Apps From a Phone

A whole web app, built from a phone.

Claude Code in the Claude app, prompt by prompt

Chapters

  1. Intro
  2. The phone is the remote
  3. Prompt one: plan first
  4. Prompt two: the first build
  5. Fine-tuning from the phone
  6. Deploy and ship
  7. Prompts that work

Who it's for

Developers and makers who want to build away from the desk, and anyone who has tried "vibe coding" and got vague results. The first chapters are plain language; the rest shows the actual prompts, replies and fixes.

Context

The sample project is lab.join.dev, the "Neon Card Studio": you type a system-design concept, a server route asks Gemini's image model for an illustration in the join.dev neon style, and the card goes into a gallery. The whole build happens in one Claude Code cloud session started from the Code tab of the Claude mobile app. The prompts, replies, commits and app screens in the video are illustrative recreations of that flow.

Architecture

ComponentRoleNotes
Claude app (phone)The remote: prompts in; replies, screenshots, diffs and preview links outCloud sessions can be started and steered from the Code tab of the Claude app, claude.ai/code, Desktop or claude --cloud
Cloud session VMRuns Claude Code, the repo, dev server, a mobile browser (Playwright WebKit) and the checksOne isolated, Anthropic-managed VM per session; idle VMs are reclaimed, pushed work survives
Cloud environmentNetwork level, setup script (npm ci, npx playwright install webkit), API credentialsSet up once at claude.ai/code, opened in the phone's browser; the app just picks it. Custom network adds cdn.playwright.dev and .vercel.app to the Trusted defaults (which already include .googleapis.com)
Plan modeClaude reads and explores but doesn't edit until you approveChosen from the mode picker when the session starts
CLAUDE.mdPlan, rules and commands for every later sessionFirst commit on the branch
/api/generateServer route: validate (≤ 80 chars), rate limit per IP + daily cap, call Gemini, store the cardThe Gemini key lives on the server (Vercel env var) and, inside the VM, as a proxy-attached API credential
web-qa subagentPlaywright's iPhone device profile (mobile WebKit, touch, phone width); reads its own screenshotsThe automated check: sends screenshots back into the chat and gates the preview pipeline
Phone's browserWhere you preview and test every buildTap the preview link, type a concept, tap Generate; it catches what emulation misses (a real screen width, touch instead of hover, the deployed environment)
GitHub + VercelBranch pushes build previewsNo deploy token in the VM; previews sit behind Vercel Authentication, so QA sends the protection-bypass header

Request flow

  1. Prompt one, plan mode: "Build lab.join.dev: I type a system design concept and Gemini draws a neon card. Plan it first. Don't write code yet." → a five-point plan.
  2. Tighten the plan: "Also cap requests per visitor and per day, and keep concepts under 80 characters." → rate limit, daily cap, 400 on long concepts. Approve.
  3. Prompt two: "Go ahead and build it. Send me a screenshot when the page works." → build, check, push, a 390 px screenshot of a working but plain v1.
  4. Fine-tunes, one prompt each: v2 "Use the join.dev look: navy grid, neon cyan, a lime button." v3 "Add style chips, and a shimmer while it generates." v4 "Add a gallery of saved cards and a small tilt on hover." One commit each.
  5. Test in the phone's browser: open the preview link, type "rate limiting", tap Generate, watch the shimmer. The card runs off the right edge. "On my phone the card is cut off on the right. Fix it and send 390 px screenshots." → reproduced with web-qa at 390 px, fixed width → max-width: 100%, after screenshot.
  6. Retest on the phone: the card fits, but tapping it does nothing: the tilt was hover-only, and a phone has no hover. "There's no hover on a phone. Tilt the card when it's tapped instead." → hover kept for mouse, tap tilt for touch; reload, tap, it tilts.
  7. Test the deployed preview: open the newest preview link and Generate → "Something went wrong (500)". "Why does Generate fail on the preview but not in dev?" → GEMINI_API_KEY isn't set for Vercel previews, and the agent can't see Vercel settings. You add it in the Vercel dashboard; "Added the key in Vercel. Redeploy and test the preview again." → green pipeline.
  8. Retest what failed, on the phone: Generate "kv cache" on the new preview → the card comes back. Then review the diff in the app, approve, PR ready, and share the preview link: anyone can test it in their own phone's browser.

The failure modes

  1. The blocked network. Trusted blocks Playwright's browser download and *.vercel.app, so the setup fails and the agent can't test its own previews.
  2. What only a real phone shows. The card overflows the screen, and the hover-only tilt never fires on touch; both were caught by testing in the phone's browser (a screenshot can't show a missing tap), then reproduced in web-qa's mobile browser and fixed.
  3. It works in dev. The deployed preview has no GEMINI_API_KEY: every page loads but Generate returns 500, found by generating a card on the preview from the phone.
  4. (Underlying) Vague prompts. "Make it look better", "fix the bug" leave the agent guessing and produce changes you can't judge from a phone screen.

Why it happens

On a phone you can't read code or logs comfortably, so everything depends on the agent sending back small, checkable evidence; vague prompts and missing environment boundaries (network, secrets, deploy target) are where that evidence goes wrong.

The fix

  • Set the environment up once with least-privilege egress and keys as API credentials.
  • Plan first in plan mode, tighten the plan, then approve.
  • One change per prompt, one commit each, pushed every time.
  • Ask for proof: screenshots at named widths, check output, the preview URL.
  • Name the device or place: "on my phone", "390 px", "on the preview".
  • Ask why before fixing when something fails in one place but not another.
  • Test in the phone's browser on every preview link: type, tap, look; touch and real screen widths behave differently from a desktop check.
  • QA the deployed preview, not just dev, and retest exactly what failed after a fix; keys never go in prompts.

Trade-offs

  • The phone is great for steering and reviewing small diffs, poor for large refactors; --teleport moves the session to a terminal when needed.
  • Screenshots per change cost a little time and context; they're what make phone review possible.
  • One change per prompt means more round trips, but each is cheap to undo.
  • API credentials are Pro/Max only for now; on Team/Enterprise keys live in environment variables, which anyone using the environment can read.

Numbers worth knowing

  • Setup script cached only if it finishes in roughly 5 minutes (Claude Code docs).
  • Illustrative only: the app, prompts, replies, commit hashes, the 80-character limit, "12 cards", the preview URL and every app screen.

Six prompt habits

Plan first

Plan mode, approve, then build

One change

One fine-tune per prompt, one commit each

Ask for proof

Screenshots, check output, the preview URL

Name the device

390 px, the preview, not just dev

Why, then fix

Ask for the cause before the fix

Keys stay out

Credentials in the host, never in a prompt

Recap: before and after

Vague

  • Make it look better
  • Fix the bug
  • Build the whole app
  • Is it done?
vs

Specific

  • Use the join.dev look: navy, cyan, lime
  • The card is cut off at 390 px
  • Plan it first. Don't write code yet.
  • Send phone-size screenshots

Name the change, the place, and the proof you want back.

Sources

Coming next in the series: Agent sandboxing: how isolation really works

Found this useful?

Subscribe for the next episode, or share it with the person who owns this part of your stack.

Keep going