Cinematic deep dive Agents in Production 6 min Failure mode + fix

Building Memory for AI Agents

Your agent remembered the wrong version of you.

What to remember, when to retrieve it, and how to forget

Chapters

  1. Intro
  2. Two kinds of memory
  3. Write a useful memory
  4. Recall at the right moment
  5. When memory goes stale
  6. Update, forget, verify

Who it's for

Engineers building assistants that serve the same user across many conversations, especially support, productivity and coding agents that should adapt without treating every past sentence as permanent truth.

Context

Conversation history is thread-scoped working state. Long-term memory persists across threads. A larger context window alone does not decide which facts deserve persistence, which user owns them, or whether a later correction supersedes them. This episode follows a fictional user's travel-planning assistant. The user first prefers morning flights, later changes to evening flights, and expects the next session to honor the new preference. All story details are illustrative.

Architecture

ComponentRoleNotes
Thread checkpointerResume the current conversationShort-term state, keyed by thread ID
Candidate extractorPropose durable facts from a turnA candidate is not yet trusted memory
ReconcilerAdd, update, supersede or reject candidatesCompare with existing facts and dated evidence
Scoped memory storePersist cross-session recordsNamespace by user or organization; stable record IDs
RetrieverFind relevant memory for this taskFilter by owner and status before ranking
Context builderPlace selected facts into the promptInclude source and freshness; keep a small budget
Memory controlsInspect, correct and delete recordsUser-facing controls and cascading index deletion

Request flow

  1. A user says, “I usually prefer morning flights.” The assistant responds and records the conversation in thread state.
  2. An extractor proposes a narrowly scoped preference. The reconciler checks the existing record, provenance and whether it is appropriate to retain. The store writes a record under that user's namespace.
  3. In a later thread, the request “Find a flight to Berlin” triggers a user-scoped search. A relevant, active preference is selected, with its source, and placed in the prompt. The agent can use it while still asking when the choice matters.
  4. The user later says, “Evening flights work better now.” The reconciler supersedes the morning preference and records the newer evidence. The old value is excluded from active retrieval.
  5. If the user deletes the preference, the record and searchable index entry are removed. The assistant no longer reintroduces it from memory.

The failure modes

The missing memory. A new thread has a new state key, so an assistant with only a thread checkpointer does not see the earlier preference. A long transcript kept in one prompt can also fail to surface the relevant fact when needed. The result is a generic answer or repeated questions.

The stale memory. An append-only list stores both “morning flights” and “evening flights.” Similarity search can return the old fact, and a model may quietly plan around it even when it also sees a correction. Recent work on stale agent memory evaluates this exact supersession problem. A retrieved memory needs current status, timestamp and source, not just a similar sentence.

Why it happens

Many systems conflate chat history, persistent memory and retrieval. They write too much, treat extraction as truth, and have no update or deletion path. Similarity ranking alone cannot determine whether an old preference is still valid.

The fix

Keep thread state and cross-session memory separate. Save only durable, useful candidates. Scope every record to the right owner, and keep a stable ID, type, source, timestamp and active/superseded status. Reconcile each new candidate against current records; do not append conflicting preferences as independent active truths. At read time, filter by owner and status, then retrieve for relevance and place only a few cited facts in context. Offer inspect, correct and delete operations. Evaluate both recall of useful facts and rejection of obsolete ones.

LangGraph's checkpointer/store split and its memory overview provide concrete examples of thread state, namespaces, JSON memories, hot-path and background writes. The record schema and reconciliation policy in the episode are a proposed application design, not a prescribed LangGraph API.

Trade-offs

  • Hot-path reconciliation makes the next request accurate sooner but adds latency; background extraction keeps responses fast but can leave a short freshness gap.
  • Narrow facts are easier to update and delete than a single biography, but retrieval and conflict resolution become more complex.
  • Aggressive filtering reduces irrelevant or sensitive memories but can miss implicit preferences. Evaluate real tasks, not only direct fact questions.
  • Source retention supports correction and audit, but increases storage and privacy obligations. Set retention rules for the product.

Numbers worth knowing

  • No performance percentages are claimed. Counts and times in the animation are illustrative.
  • LoCoMo studies long-term dialogue over many sessions; STALE studies whether agents reject invalidated memories. Their results motivate separate recall and freshness checks, not a universal performance promise.

The read path in code

memory.py
def context_for(user_id, task):
    matches = store.search(
        namespace=(user_id, 'memory'),
        filter={'status': 'active'},
        query=task, limit=3)
    return [
        {'fact': m.fact, 'source': m.source,
         'updated_at': m.updated_at}
        for m in matches
    ]

What earns a place in memory

Useful

Likely to change a future answer

Specific

One fact with a clear owner and scope

Traceable

Keep source and update time

Appropriate

Apply product retention rules

Recap: before and after

Append only

  • Morning and evening both active
  • Old result can rank first
  • Deletion leaves index copy
vs

Reconcile

  • New evidence supersedes old
  • Retrieve active records
  • Delete record and index entry

Freshness is a data rule before it is a model task

Sources

Coming next in the series: When RAG Leaks Private Documents

Found this useful?

Subscribe for the next episode, or share it with the person who owns this part of your stack.

Keep going