Cinematic deep dive Agents in Production 5 min Failure mode + fix

Why AI Agents Act Twice

The tool succeeded. Your agent did it again.

Durable execution, idempotency, and the missing reply

Video premiering soon on @AI.JoinDev Subscribe to get it first

Chapters

  1. One request, two effects
  2. Separate decisions from effects
  3. The missing acknowledgement
  4. Recover the recorded decisions
  5. Give the action a stable identity
  6. Handle the uncertain cases

Who it's for

Engineers building agents that create records, dispatch shipments, or call other tools with external effects.

Context

An illustrative shipping agent sends an authorized request. The provider commits it, but an acknowledgement is lost. A worker timeout or crash leaves the caller unable to distinguish failure from success. This is a distributed execution problem even when the model proposed the correct action.

Architecture

ComponentRoleNotes
Agent loopProposes and orchestrates actionsModel output is data, not permission
Authorization boundaryValidates intent and argumentsSaved intent includes tenant and operation identity
Durable workflowPreserves execution progressRecorded model results are reused on replay
Tool activityAttempts the external operationUnrecorded completion can lead to another attempt
ReceiverEnforces idempotencySame key and arguments must not repeat the effect within its contract
Action ledgerStores intent, attempts, receipts, uncertaintyA local receipt alone cannot make a remote effect atomic

Request flow

  1. Validate and authorize an action; persist its identity and immutable arguments before dispatch.
  2. Schedule the tool activity, carrying the persisted identity into every attempt.
  3. At the receiver, atomically couple deduplication with a local effect, or use the remote provider's documented idempotency interface.
  4. Persist the returned receipt and record workflow completion.
  5. After an ambiguous timeout, reuse the same identity only within the provider's guarantees; otherwise reconcile by reliable business reference and escalate unresolved outcomes.

The failure modes

  • Lost reply: the service succeeds but the acknowledgement disappears. A naive retry creates a second shipment.
  • Crash before acknowledgement: the worker succeeds remotely and crashes before recording its receipt. A replacement worker repeats the action because its local history has no proof of completion.

Why it happens

The receiver's external effect and the caller's completion record are separate commits. Neither a longer timeout nor saving chat history closes that gap.

The fix

Use durable execution for recorded progress and receiver-enforced idempotency for repeated effects. Temporal's workflow replay reconstructs recorded state; completed activity results in history are reused, but an activity whose completion was not recorded may execute again. Model calls belong in activities, not in deterministic workflow logic.

An operation ID represents one authorized business intent, not an attempt. Scope it by tenant and action, preserve exact arguments, reject conflicting reuse, and issue a fresh identity for a genuinely new action. A shipment provider with an idempotency contract is illustrative; the pseudocode is not a real shipping SDK.

Stripe supplies a concrete API contract: key reuse compares parameters, cached results include errors, and retention is finite. Its documentation permits pruning keys after at least 24 hours; do not assume an old key still deduplicates. Concurrent requests or pre-execution validation failures have additional rules. Check the actual provider before choosing retry behavior.

A transactional outbox solves atomic local state plus dispatch intent. Its relay may deliver duplicates, so consumers still need idempotency. When an external service offers neither deduplication nor a reliable outcome lookup, preserve the unknown state and require reconciliation rather than claiming exactly-once external execution.

Trade-offs

  • Durable state, receipt retention, and reconciliation add storage and operational complexity.
  • Finite deduplication retention bounds safe automatic retry windows.
  • Granular activities improve recovery visibility but increase history volume.
  • Protect sensitive stored arguments and receipts; access control remains separate from retry correctness.
  • Backoff with jitter and bounded attempts avoid retry amplification but cannot repair an ambiguous effect by themselves.

Numbers worth knowing

No performance benchmarks or measured latency claims are presented. The shipment scenario and pseudocode are illustrative. The 24-hour Stripe retention detail above is sourced and deliberately omitted from narration to keep the lesson provider-independent.

Identity belongs to the saved intent

tool-boundary.pseudo
intent = load_authorized_intent(operation_id)
key = intent.tenant + ":ship:" + intent.id
args = intent.validated_arguments

# Provider must enforce idempotency.
receipt = shipping.create(
    args, idempotency_key=key
)

save_receipt(intent.id, receipt)

Durable execution remembers progress

Persist progress

Record completed steps and their results.

Record model output

Replay facts instead of asking again.

Retry unfinished work

External activities may execute again.

Recap: before and after

Blind retries

  • New key on every attempt
  • Assume timeout means failure
  • Retry unknown outcomes forever
vs

Explicit action state

  • Stable identity and arguments
  • Receipt or reconciliation
  • Bounded retries and escalation

Unknown is a real state.

Sources

Coming next in the series: Agent memory: what should survive?

Found this useful?

Subscribe for the next episode, or share it with the person who owns this part of your stack.

Keep going