Cinematic Short More system design 47 sec Failure mode + fix

Approval Fatigue

You tapped Allow 41 times. One was rm -rf.

Video premiering soon on @AI.JoinDev Subscribe to get it first

Why it happens

Constant prompts turn Allow into a reflex. The dangerous one gets the same tap as forty safe ones.

The fix

Allow routine commands, deny dangerous ones, sandbox files and network. Rare prompts get read.

Context

A developer runs a coding agent (Claude Code in the video) against a local repo and approves its actions from a phone while away from the desk. The agent runs in a manual "ask before acting" setup with no allow, deny or sandbox rules, so every shell command and every file edit raises a permission prompt. A single bug fix can easily produce dozens of prompts (the count of 41 in the video is illustrative). The stakes are the developer's machine: the agent runs with their user account, so a destructive command can reach far beyond the project.

Architecture

ComponentRoleNotes
DeveloperGives the task, answers promptsOften on a phone, between other work
Coding agent (Claude Code)Plans and runs commands and editsRuns with the developer's user permissions
Approval promptThe only control: a human "Allow / Deny"Same look and weight for npm test and rm -rf
Shell + filesystemWhere commands take effectHome folder, SSH keys, other repos are all reachable

Request flow

  1. The developer asks the agent to fix a bug.
  2. The agent wants to run npm test: a prompt arrives, the developer reads it and taps Allow.
  3. Edits to src/app.ts, src/auth.ts, a lint run, more tests: each one prompts, each one is safe, each one gets Allow.
  4. Around the fortieth prompt the developer stops reading and taps Allow on sight.
  5. The next request is rm -rf ~/ dist (a stray space turns "delete ~/dist" into "delete the home folder, then dist").
  6. It is allowed like the others, and the home folder is wiped.

The failure mode

The villain is the approval system itself. A human gate is only as good as the attention behind it, and a stream of low-risk prompts drains that attention. This is the same mechanism attackers exploit in MFA fatigue ("push bombing"): send enough push requests and someone eventually taps Approve out of habit or annoyance. With agents nobody even has to attack: the agent's own routine work produces the flood, and the dangerous action can come from a model mistake (a wrong path) or from a prompt injection hidden in a file, issue or web page the agent read (curl … | sh, pushing a secret). It is easy to miss because every prompt looks like a safety feature; nobody measures whether they are read.

Why it happens

Every action, routine or risky, goes through the same human prompt, so the prompt carries no signal. Frequent identical prompts train a reflex, and reflexes don't read commands.

The fix

Layered permissions, so the human only sees what genuinely needs a human:

  1. Allow routine, low-risk actions so they never prompt: tests, builds, lint, reads inside the project (in Claude Code: permissions.allow, e.g. Bash(npm run ), Bash(npm test )).
  2. Deny dangerous patterns outright, with no prompt to click through (e.g. Bash(rm -rf ), Bash(curl ), Read(./.env), Read(./secrets/**)). Claude Code evaluates rules in the order deny, then ask, then allow, so an allow rule can't carve an exception out of a deny.
  3. Sandbox what commands can touch: writes limited to the project directory and network limited to approved hosts (OS-level filesystem and network isolation). Claude Code's docs warn that Bash deny rules match command text and are not a security boundary on their own (/usr/bin/curl or sh -c '…' slip past Bash(curl *)), so the sandbox is the real boundary and the rules are the ergonomics.
  4. Escalate only genuinely risky actions outside the sandbox (deleting outside the project, pushing, deploying, new network hosts) with a clear, rare prompt. Rare prompts get read.

Trade-offs

  • Writing and maintaining allow/deny rules takes effort, and over-broad allow rules (Bash(*)) recreate the problem.
  • Pattern rules are brittle: they match the command text Claude writes, not the program; rely on the sandbox for safety.
  • A sandbox can break legitimate workflows (tools that write to ~/.cache, package installs needing new hosts) until configured.
  • Fully automatic modes (bypassPermissions) remove the fatigue by removing the gate; Anthropic says to use them only in isolated containers or VMs. Classifier-based approval (Claude Code's auto mode) is a middle ground, not a guarantee.

Numbers worth knowing

  • Anthropic reports that sandboxing reduced permission prompts by 84% in its internal usage of Claude Code (Anthropic engineering, "Beyond permission prompts: making Claude Code more secure and autonomous", Oct 2025).
  • "41 prompts" and "forty safe prompts" in the video are illustrative.
  • CISA's fact sheet "Implementing Number Matching in MFA Applications" (Oct 2022) recommends number matching to mitigate MFA fatigue: make the approval require reading, not just tapping.

Sources

  • Claude Code docs, "Configure permissions" (code.claude.com/docs/en/permissions): permission modes, allow/ask/deny rules,

  • Claude Code docs, "Sandboxing" (code.claude.com/docs/en/sandboxing).

  • Anthropic engineering, "Beyond permission prompts: making Claude Code more secure and autonomous" (Oct 20, 2025):

  • CISA, "Implementing Number Matching in MFA Applications" fact sheet (Oct 2022): MFA fatigue / push bombing.

Coming next in the series: Agent Sandboxing

Found this useful?

Subscribe for the next episode, or share it with the person who owns this part of your stack.

Keep going