Deep dive Agents in Production 4 min Failure mode + fix

Autonomous Coding Agents

The hard part is not writing code. It is proving the change is safe.

How a ticket becomes a reviewable pull request

Chapters

  1. The real job
  2. A bounded workspace
  3. The agent loop
  4. One task, step by step
  5. Villain one: missing context
  6. Villain two: green tests
  7. Production guardrails
  8. The reviewable handoff
  9. Recap

Who it's for

Engineering leaders and developers introducing coding agents to a production repository without turning autonomy into unchecked deployment.

Context

The workspace sits between an issue tracker and a pull request. It gives a model the task, repository instructions, a sandbox, constrained tools, and explicit quality gates. The goal is not a magical one-shot patch; it is a traceable engineering loop.

Architecture

ComponentRoleNotes
Task intakeTurns issue into scoped objectiveIncludes acceptance criteria and risk level
Context builderReads repository instructions and relevant codeKeeps evidence linked to files and commands
PlannerProposes small, ordered changesA plan is revisable, not authority to deploy
Sandbox + toolsSearch, edit, test, browser and GitLeast privilege and isolated execution
EvaluatorRuns deterministic checks and reviews diffsGates the next iteration
PR handoffSummarizes evidence for a human reviewerHuman owns merge and release

Request flow

  1. A developer assigns a well-scoped issue with acceptance criteria.
  2. The agent reads repository instructions, searches for related code, then writes a plan.
  3. It edits a small slice, runs focused tests, and observes the results.
  4. It repeats only when evidence says the task is incomplete.
  5. Lint, tests, policy and diff review form gates before a pull request is prepared.
  6. A human reviewer decides whether to merge.

The failure modes

The context trap

An agent starts coding from the ticket alone, misses local conventions or a nearby implementation, and makes a plausible patch in the wrong layer. The patch can compile while violating an unstated contract.

The green-test illusion

An agent treats one passing command as proof. The changed path is untested, snapshots bless a regression, or tests never ran against the intended files.

Why it happens

Models act on the context and tools they are given. Vague tasks, broad permissions, weak checks, and opaque execution remove the feedback that ordinary engineering relies on.

The fix

Use a bounded agent loop: scoped task → repository-aware plan → least-privilege sandbox tools → small diffs → layered verification → human review. Treat every tool result as new evidence, not as a ceremony. Record the plan, changed files, commands, results, and unresolved risks in the PR.

Trade-offs

  • More context and verification improve reliability but cost time and tokens.
  • Narrow permissions reduce blast radius but can require human escalation.
  • Small commits are easier to review but may take more loop iterations.
  • Passing tests are necessary evidence, not a guarantee of correctness.

Numbers worth knowing

  • No universal throughput or correctness numbers are claimed here; any timings and percentages shown are illustrative workflow budgets.

Policy before execution

agent-policy.ts
const policy = {
  workspace: 'isolated',
  writablePaths: ['app/', 'tests/'],
  commands: ['npm test', 'npm run lint'],
  network: 'deny',
  requireReview: true,
  allowMerge: false,
};

What the reviewer receives

Intent

Issue, acceptance criteria, and the plan the agent followed.

Context

Files inspected and the repository rules that shaped the patch.

Diff

Small, reversible changes with a concise rationale.

Evidence

Commands, test results, checks, and known limitations.

Decision

A human decides to request changes, merge, or reject.

Recap: before and after

One-shot code bot

  • Vague prompt
  • Broad permissions
  • Large opaque patch
  • One test command
  • Implicit deployment risk
vs

Coding agent workspace

  • Scoped acceptance criteria
  • Least-privilege sandbox
  • Small evidence-led edits
  • Layered verification
  • Human merge decision

Autonomy earns trust through a visible feedback loop.

Sources

  • OpenAI, “Unrolling the Codex agent loop.”

  • OpenAI, “Running Codex safely at OpenAI.”

  • Anthropic, “Demystifying evals for AI agents.”

  • GitHub Docs, “About third-party coding agents.”

Coming next in the series: Prompt Injection Through Tools

Found this useful?

Subscribe for the next episode, or share it with the person who owns this part of your stack.

Keep going