Building Memory for AI Agents
Your agent remembered the wrong version of you.
What to remember, when to retrieve it, and how to forget
Chapters
- Intro
- Two kinds of memory
- Write a useful memory
- Recall at the right moment
- When memory goes stale
- Update, forget, verify
Who it's for
Engineers building assistants that serve the same user across many conversations, especially support, productivity and coding agents that should adapt without treating every past sentence as permanent truth.
Context
Conversation history is thread-scoped working state. Long-term memory persists across threads. A larger context window alone does not decide which facts deserve persistence, which user owns them, or whether a later correction supersedes them. This episode follows a fictional user's travel-planning assistant. The user first prefers morning flights, later changes to evening flights, and expects the next session to honor the new preference. All story details are illustrative.
Architecture
| Component | Role | Notes |
|---|---|---|
| Thread checkpointer | Resume the current conversation | Short-term state, keyed by thread ID |
| Candidate extractor | Propose durable facts from a turn | A candidate is not yet trusted memory |
| Reconciler | Add, update, supersede or reject candidates | Compare with existing facts and dated evidence |
| Scoped memory store | Persist cross-session records | Namespace by user or organization; stable record IDs |
| Retriever | Find relevant memory for this task | Filter by owner and status before ranking |
| Context builder | Place selected facts into the prompt | Include source and freshness; keep a small budget |
| Memory controls | Inspect, correct and delete records | User-facing controls and cascading index deletion |
Request flow
- A user says, “I usually prefer morning flights.” The assistant responds and records the conversation in thread state.
- An extractor proposes a narrowly scoped preference. The reconciler checks the existing record, provenance and whether it is appropriate to retain. The store writes a record under that user's namespace.
- In a later thread, the request “Find a flight to Berlin” triggers a user-scoped search. A relevant, active preference is selected, with its source, and placed in the prompt. The agent can use it while still asking when the choice matters.
- The user later says, “Evening flights work better now.” The reconciler supersedes the morning preference and records the newer evidence. The old value is excluded from active retrieval.
- If the user deletes the preference, the record and searchable index entry are removed. The assistant no longer reintroduces it from memory.
The failure modes
The missing memory. A new thread has a new state key, so an assistant with only a thread checkpointer does not see the earlier preference. A long transcript kept in one prompt can also fail to surface the relevant fact when needed. The result is a generic answer or repeated questions.
The stale memory. An append-only list stores both “morning flights” and “evening flights.” Similarity search can return the old fact, and a model may quietly plan around it even when it also sees a correction. Recent work on stale agent memory evaluates this exact supersession problem. A retrieved memory needs current status, timestamp and source, not just a similar sentence.
Why it happens
Many systems conflate chat history, persistent memory and retrieval. They write too much, treat extraction as truth, and have no update or deletion path. Similarity ranking alone cannot determine whether an old preference is still valid.
The fix
Keep thread state and cross-session memory separate. Save only durable, useful candidates. Scope every record to the right owner, and keep a stable ID, type, source, timestamp and active/superseded status. Reconcile each new candidate against current records; do not append conflicting preferences as independent active truths. At read time, filter by owner and status, then retrieve for relevance and place only a few cited facts in context. Offer inspect, correct and delete operations. Evaluate both recall of useful facts and rejection of obsolete ones.
LangGraph's checkpointer/store split and its memory overview provide concrete examples of thread state, namespaces, JSON memories, hot-path and background writes. The record schema and reconciliation policy in the episode are a proposed application design, not a prescribed LangGraph API.
Trade-offs
- Hot-path reconciliation makes the next request accurate sooner but adds latency; background extraction keeps responses fast but can leave a short freshness gap.
- Narrow facts are easier to update and delete than a single biography, but retrieval and conflict resolution become more complex.
- Aggressive filtering reduces irrelevant or sensitive memories but can miss implicit preferences. Evaluate real tasks, not only direct fact questions.
- Source retention supports correction and audit, but increases storage and privacy obligations. Set retention rules for the product.
Numbers worth knowing
- No performance percentages are claimed. Counts and times in the animation are illustrative.
- LoCoMo studies long-term dialogue over many sessions; STALE studies whether agents reject invalidated memories. Their results motivate separate recall and freshness checks, not a universal performance promise.
The read path in code
def context_for(user_id, task):
matches = store.search(
namespace=(user_id, 'memory'),
filter={'status': 'active'},
query=task, limit=3)
return [
{'fact': m.fact, 'source': m.source,
'updated_at': m.updated_at}
for m in matches
]
What earns a place in memory
Useful
Likely to change a future answer
Specific
One fact with a clear owner and scope
Traceable
Keep source and update time
Appropriate
Apply product retention rules
Recap: before and after
Append only
- Morning and evening both active
- Old result can rank first
- Deletion leaves index copy
Reconcile
- New evidence supersedes old
- Retrieve active records
- Delete record and index entry
Freshness is a data rule before it is a model task
Sources
LangChain, Memory overview: thread versus cross-session memory, memory types, update timing, namespaces and JSON records.
LangGraph, Memory guide: checkpointers and long-term stores.
Maharana et al., LoCoMo (2024): long-term conversational memory evaluation across sessions.
Chao et al., STALE (2026): invalidated memories and implicit conflicts.
Chang and Chen, LOCOMO-CONV (2026): implicit conversational memory retrieval beyond direct fact questions.
Coming next in the series: When RAG Leaks Private Documents