Home / Blog / Context engineering and memory / Agent Memory Architecture: A Taxonomy and Three…

Context engineering and memory

Agent Memory Architecture: A Taxonomy and Three Designs

Agent memory architecture in one page: four kinds of memory ranked by write risk, three designs compared, and a write policy to copy. Pick your design.

By Enrique Gutiérrez · Published · 23 min read

An agent memory architecture is the set of decisions about what an agent keeps beyond a single model call: which kinds of memory it has, where each piece waits, what moves it back into the context window, and who may write it. Those decisions are yours to make, because the model keeps nothing between calls.

I wrote this for the engineer who owns the evals and has been asked to add memory to an AI agent. By the end you should be able to name the kinds your agent needs, pick one of three designs from two inputs, and paste a write policy into the design note.

What is an agent memory architecture?

An agent memory architecture answers one question with two parts, which Chapter 9 of AI Agents, Engineered puts this way: “Every memory architecture in the field is an answer to the same two-part question: where does each piece wait, and what moves it.”

The book’s definition of memory is plain about who does the work. “Memory is a system somebody engineered: a set of standing decisions about what to keep, where to keep it, and how to bring the right piece back onto the desk at the right moment.” The desk is the book’s name for the context window.

The practical steps of how to give an AI agent memory without making it worse are a separate job. I also leave out three neighboring subjects: the craft of context compaction, the diagnosis of context rot, and the choice between retrieval and a long window for outside knowledge. Whether a run survives a crash is a reliability question that the book gives to Chapter 18.

What kinds of memory does an agent have?

An agent has four kinds of memory in the book’s taxonomy: short-term memory, which the field also calls working memory, and three long-term kinds named episodic, semantic and procedural. Chapter 9 makes the first cut by lifetime and the second by content. The two lifetimes get one sentence: “Working state belongs to the run; durable knowledge belongs to the relationship.”

Short-term memory is the live window itself, and the glossary equates it with the message history. Long-term memory is everything kept outside the window and carried across sessions. The book gives each long-term kind a piece of office furniture.

Kind (book’s name) What it holds Lifetime Where it is stored How it is retrieved Cost of a bad write
Short-term, or working Standing instructions, the request, the transcript, tool results, text the agent wrote to itself One run or conversation; compaction, truncation or a fresh session ends it The context window No retrieval: the model re-reads all of it on every call Bends the rest of the run; a fresh window clears it †
Episodic (the logbook) What happened: past conversations, past tool-call sequences Across sessions; append-only; old episodes may fade once consolidated A store, queried by relevance Search, blended with recency Low: one record among several, inert until a search surfaces it
Semantic (the filing cabinet) What is known, stripped of the episode that taught it Across sessions, until updated, reconciled or expired A store, as small fact documents or one profile; a notes file at small scale † Search, filtered by owner first Medium: cited by every answer that retrieves it, until reconciliation or expiry
Procedural (the procedures manual) How things are done here: standing instructions, playbooks, learned rules Across sessions, until someone edits it A standing instructions file, always loaded No search: loaded into every session High: changes behavior in every session and every task

Cells marked † are my additions. Everything else restates Chapter 9, including the ranking in the last column, although the words low, medium and high are my shorthand for its three paragraphs. The book ranks only the three long-term kinds. I added the working-memory cell from its remark that a poisoned window is better restarted than repaired.

The furniture suggests three ways of reading, and the book corrects that: there is one, retrieval. Chapter 9 says a memory store “is simply one more corpus, and reading from memory is retrieval-augmented generation over your own history”. The comparison of RAG and agents is a separate question about who decides when to retrieve.

Where do other taxonomies agree and differ?

The outside literature agrees with the book on the four names and differs on what two of them mean. The names come from Sumers and colleagues (arXiv:2309.02427 v3), the paper the book credits. That paper defines working memory more broadly than the window: “CoALA’s notion of working memory is more general: it is a data structure that persists across LLM calls.”

Procedural memory differs too. Sumers and colleagues count “two forms of procedural memory: implicit knowledge stored in the LLM weights, and explicit knowledge written in the agent’s code”. The book uses the narrower sense, the instructions and playbooks a session loads, and so does this post.

Other sources count differently. The system that introduced paging to this field (Packer and colleagues, arXiv:2310.08560 v2) sorts memory by location and describes its technique as borrowed from operating-system memory hierarchies, with “data movement between fast and slow memory”. A 2026 survey by Pengfei Du (arXiv:2603.07670 v1) keeps the four names and adds two more axes, which it calls “representational substrate” and “control policy”.

Why does write risk order the kinds?

Write risk orders the kinds because the three long-term kinds differ in how a bad entry reaches the model. An episode waits for a search to find it, a fact is cited whenever it is retrieved, and a rule is loaded every time. Chapter 9 says “the three kinds of memory are not equally dangerous to write” and ends the argument in two sentences: “A wrong fact costs you an answer. A wrong rule costs you the agent.”

Sumers and colleagues make the same point. They write that learning by writing to procedural memory “is significantly riskier than writing to episodic or semantic memory, as it can easily introduce bugs or allow an agent to subvert its designers’ intentions”. Their sentence concerns agents that rewrite their own code, and it applies as well to one that appends lessons to its instructions file.

In an agent memory architecture the ranking tells you where to spend review effort. Chapter 9 asks that procedural edits be “reviewed before they land, versioned so they can be rolled back, and small enough to be understood”. Semantic writes need a reconciliation rule and an expiry. Episodic writes need little more than a timestamp and an owner.

What are the three designs?

The three designs are a notes file the agent reads and writes, a store with retrieval, and a chain of short runs that pass notes forward. Grouping them as three architectures is this post’s own framing. Chapter 9 lists a file, stores and paging between tiers under its heading on architectures, and gives the chain of short runs later, as its advice for one long run.

Paging, the book’s third item, describes movement between three tiers: hot (always loaded), warm (loaded when a task of that kind begins) and cold (retrieved on demand or never). I treat paging as the loading policy inside each design. In a file design the hot tier is the part read at every start; in a store design the cold tier is the store.

The designs also compose: a support agent can keep a store for user facts and still run a long case as a chain of short runs.

What is a notes file, and how does it fail?

A notes file is a plain document that the agent reads at the start of a session and updates as it works. Chapter 9 opens its list with it: “The simplest durable memory is a file.” Its examples are a progress note, a plan the agent ticks off, and a notes document reread after every reset. The book calls this the baseline, since a person can open the file, find the misleading sentence and fix it.

Coding agents already ship this category. Two labeled examples, both read on 2026-10-07: the Claude Code memory documentation describes instruction files a person writes and “notes Claude writes itself based on your corrections and preferences”, both loaded at the start of every conversation. The AGENTS.md site describes its convention as “a README for agents”. I name them as examples of the category and recommend neither.

The cost is rent on every call: whatever loads at the start occupies the window all session, so the book’s rule for the hot tier applies: “keep it ruthlessly small”. The context window budget planner has a line for this component.

A notes file fails by going stale. One practitioner reported in March 2026 that with file-based memory “after about 3 weeks the recall quality drops noticeably”. Another wrote in October 2026: “It’s like the LLMs can’t just DELETE something, they always amend.” A file has no search to hide an old line, so every stale sentence is read again each session.

What is a store plus retrieval, and how does it fail?

A store plus retrieval is a database the agent reads and writes through tools, with a search step that brings back a few records per call. Chapter 9 describes stores as “databases the agent reads and writes through tools”, usually a vector store beside plainer ones. It reports that production designs commonly file each memory under a namespace and a key, and states the trade: “The store’s advantage over the file is scale and search; its cost is that nobody can proofread a million embeddings.”

Managed stores are a second product category. Two labeled examples, described from the abstracts of their own papers: Mem0 (arXiv:2504.19413, April 2025) presents an architecture for “dynamically extracting, consolidating, and retrieving salient information from ongoing conversations”. Zep (arXiv:2501.13956, January 2025) presents “a temporally-aware knowledge graph engine”. Both papers report benchmark gains measured by their own authors, so I print none of the figures.

The cost is a second system to run: an extraction step, an index, a search on every call, and upkeep. A store you buy still leaves the write policy and the reconciliation rule to you.

A store fails in three ways the book names: stale facts, contradictions and poisoning. A contradiction appears wherever the write path lacks a reconciliation rule, and then “retrieval, blameless, returns both”. One practitioner described the result in March 2026: “Over time the agent is recalling contradictory facts with equal confidence.” Junk is the fourth, and the book calls it the store-everything trap.

What is a chain of short runs, and how does it fail?

A chain of short runs splits one long task into several sessions, each of which starts in a fresh window seeded from notes the previous one wrote. Chapter 9 gives it as the one-line summary of its protocol for long runs: “a healthy long run is a chain of short runs joined by good notes”. The protocol’s steps come before that line: trim tool results, offload to files, fetch late, compact, prune, and start fresh.

Files and stores carry durable knowledge across sessions. A chain carries working state across the death of a window, and the notes it passes are working state written outward. The book’s reason for writing them early is short: “files survive what summaries forget”.

The cost is a few read and write steps per run, plus the discipline of ending each run at a clean boundary.

A chain fails when the note omits the one live constraint, or when the next run never rereads it. It also fails quietly when notes from a finished task are kept and treated as facts, so run notes should expire with the task unless something promotes them on purpose.

How do the three designs compare?

The three designs differ most in how you debug them and in what goes wrong first. The table is my summary. Its scale and debugging rows follow Chapter 9, and its last row points to the tests further down.

Notes file Store plus retrieval Chain of short runs
Carries Durable knowledge, small Durable knowledge, large or per owner Working state for one long task
What moves the pieces Loaded at session start A search on each call A fresh window reads the notes
Scales to What a session can load whole Many records and many owners Task length; each run stays short
How you debug it Open the file and read it Query the store; inspect records and scores Read the note that crossed the boundary
Fails first by Stale lines, never deleted Contradictions, stale facts, junk A constraint missing from the note
Costs Tokens on every call An extra system, plus upkeep A few steps per run
Measure with Forced reset, changed-fact pair Changed-fact pair, cross-owner leak Forced reset at a boundary

Which agent memory architecture fits your agent?

The agent memory architecture that fits depends on two inputs: the longest thing the agent must remember, and whether it reads content that anyone can write. This decision table is my own construction. The book supplies the designs and the rules; it has no such table.

Longest thing to remember What the agent reads The one rule that must hold Design
One session that fits the window Only material you control No store. The agent proposes lines for the instructions file; a person merges them Notes file
One session that fits the window Content anyone can write No store. A reader with no write access handles that content, and nothing it returns becomes a rule Notes file
One task longer than the window Only material you control Each run ends by writing plan, decisions and open problems; the next run starts from them Chain of short runs
One task longer than the window Content anyone can write Notes record outside content as claims with their source; run notes expire with the task Chain of short runs
Many sessions for one user or project Only material you control One file, reviewed like code; add a store only after an observed failure Notes file
Many sessions for one user or project Content anyone can write The file changes only through the checked write path; every line carries its source Notes file
Many users, each remembered across sessions Only material you control Filter by owner before scoring; supersede old facts Store plus retrieval
Many users, each remembered across sessions Content anyone can write Owner filter, plus the full write policy below; outside content never writes directly Store plus retrieval

The file rows follow the book against a published recommendation. Du’s survey names three patterns of its own: monolithic context (A), context plus a retrieval store (B), and tiered memory with learned control (C). It advises: “start with Pattern B, instrument it thoroughly, and graduate to Pattern C only when empirical data shows that learned control meaningfully improves your target workload.”

Pattern B is the store, so the survey’s default agent memory architecture is my second design. Chapter 7 starts lower. It says most tasks need only two cheap moves, trimming tool results and writing the plan to a file, and that “no paging memory system should exist before an observed failure asks for it”. I side with the book for one project or one user, because a file’s failures are visible in a diff, and with Du for many users, where owner filtering and search are requirements from the first day.

The rows escalate. Move from a file to a store when you observe one of two failures: the notes no longer fit what a session can load, or records must be separated by owner. Until then the trade is the one a practitioner described in October 2026, of things “the agents will burn lots of tokens to rediscover in every session” against the upkeep of notes.

How do you keep untrusted content out of durable memory?

You keep untrusted content out of durable memory by denying it a direct write: the component that reads it has no memory tool, and only a checked, attributed summary may be stored. This rule is my extension of the book. Chapter 9 states quarantine as a way to protect the main window, and gives the memory defense in a single clause.

Here is what the book says. On the window, it advises: “let a subagent with a minimal toolset and no privileges do the reading”. The subagent returns a summary, and the book is candid about the limit: “Containment reduces exposure; it does not neutralize it.” On memory, it says only that “filtering belongs on the write path, before persistence”. Joining the two is my step, and the book does not take it.

The reason to join them is the book’s own description of poisoning: “everything an agent writes from untrusted input is a chance for an attacker’s sentence to become your agent’s long-term belief”. Indirect prompt injection attacks, which hide instructions in content a tool returns, are how such a sentence arrives. A memory write makes the effect outlast the session.

on reading untrusted content (web page, email, upload):
    worker = fresh context; tools: read only; memory tools: none
    claim  = worker.summarize(content, brief)     # evidence
    return claim to the main agent

on a proposed durable write (kind, text, origin, owner):
    if kind is procedural:
        open a change for human review; stop
    if origin is untrusted and the check has not passed:
        keep it in run notes, labeled as a claim; stop
    trusted = origin is a trusted source
    record = {text, owner, source: origin, written_at: now,
              confirmed_at: now if trusted else none,
              status: active if trusted else claimed}
    old = store.find(owner, same subject as record)   # active or claimed
    if old exists and trusted:
        old.status = superseded; record.supersedes = old.id
    store.put(record)

What does the evidence on memory poisoning show?

The published evidence shows that poisoning works in research settings and says nothing about how often it happens in deployed systems.

Dong and colleagues (arXiv:2503.03704 v5) study an attacker with no write access, who acts “by only interacting with the agent via queries and output observations”. They test three agents and four types of victim and target pairs. Across those settings they report “a high average success rate of 98.2% for injecting malicious records into the memory, and a high average attack success rate of 76.8% in eliciting the malicious reasoning steps”. Their stated assumption is “that a shared memory bank is adopted to support the execution of all user queries”.

In one of the paper’s settings the attack success rate “drops quickly from 68.9% to 31.1% as benign data increases”. On defenses, it finds that detection prompts “that are precise lack generality, while general prompts may incur high false positives”. It adds that isolating memory per user can be circumvented “by identity disguise”. A filter and an owner scope reduce exposure, and neither closes it.

Chen and colleagues (arXiv:2407.12784 v1) call their work “a novel red teaming approach”. Their attacker can place entries in the memory or knowledge base, and on three types of agents they report “an average attack success rate higher than 80%” at “a poison rate less than 0.1%”. Both figures are success rates under attack conditions.

One disclosure concerns a shipped product. In September 2024 the security researcher Johann Rehberger demonstrated that content from a website could write a lasting memory in one consumer chat assistant. He reported that the vendor mitigated the data-leak route, and wrote: “A website or untrusted document can still invoke the memory tool to store arbitrary memories.” The OWASP announcement of December 2025 lists “ASI06 – Memory & Context Poisoning”; I read the announcement only.

What goes in the memory-write policy?

The memory-write policy names one writer per kind of memory, the fields every record carries, and the checks that run before anything persists. The policy is my own, assembled from the book’s rules and the write-path list in Du’s survey, which includes “metadata tagging (timestamp, source, task label, confidence)”.

MEMORY-WRITE POLICY: [agent name] · owner: [team] · reviewed: [date]

Kinds in use (delete the ones this agent does not keep)
  Working state      window + run notes [file]      writer: the agent
  Episodic           [log or store]                 writer: the harness, append-only
  Semantic           [file | store]                 writer: the write step, after the check
  Procedural         [instructions file]            writer: a person, by review

Trust
  Trusted sources:   [list: the user in session, the repository, ...]
  Untrusted sources: [list: web pages, inbound email, uploads, tool results from ...]
  Untrusted content is read by: [a worker with no memory tool | none is read]
  What the worker returns is evidence. It is never stored as a rule.

The check before a semantic write
  [ ] One fact about one subject, in the write step's own words
  [ ] Scoped to an owner: [user | project]
  [ ] Source and date recorded
  [ ] From an untrusted source: stored as "claimed by [source]" and
      confirmed by [the user | a trusted source] before an action depends on it
  [ ] Personal data: [redacted at write time | allowed fields: ...]

Every record carries
  owner · source · written_at · confirmed_at · status [active | claimed | superseded | expired]
  · supersedes [id]

Manage
  A confirmed fact on the same subject supersedes the old one: a status change, no second active record.
  A claim supersedes nothing until it is confirmed.
  Expiry: episodic [n days] · semantic [n days since confirmed_at] · run notes [task end]
  Deletion on request: [how, and within how long]

Read
  Filter by owner before scoring. Exclude superseded and expired records. Label claims as claims.
  Rank by relevance and recency.

Procedural changes
  The agent may propose a line. A person reviews and merges it. Never from untrusted content.

Tests on every policy change
  forced reset · changed-fact pair · cross-owner leak · planted-instruction write

The lethal trifecta audit is the companion check for actions. A memory write lets untrusted content reach a later session that may hold all three capabilities the audit counts.

How do you keep a memory store from rotting?

A memory store stays healthy when every record has a date, a source and an owner, when a confirmed fact replaces the old one, and when old records are allowed to expire. Chapter 9 asks for “timestamps on everything, and a willingness to let old records die”. It names consolidation, deduplication and expiry as the manage stage of a loop it credits to Du’s survey.

The lifecycle of one remembered thing.
Figure 9.3 The lifecycle of one remembered thing. Material is written from the desk to the store by two paths—in the hot path during the turn, or in a background pass afterward; the store is managed so it does not rot—consolidating episodes into facts, deduplicating, expiring the stale; and the right memory is read back by relevance and recency together. Manage, in accent, is the stage every design underbudgets. Reuse this diagram

Staleness is the failure the book treats as characteristic. It warns that “a confidently recalled obsolete fact is worse than no recall, because it arrives wearing the authority of the store”. The read side needs more than similarity for the same reason: “relevance alone is a mistake here, because an agent’s past is full of things that are similar to the present and no longer true”.

Supersede instead of append is my phrasing for the reconciliation rule the book says a write path needs. Two open issues in one open-source memory library, both unresolved when I read them on 2026-10-07, show why it matters. One reports that add-only extraction can store conflicting facts as separate records. The other reports a search that ranks an outdated memory above its update.

Whether a superseded fact should be kept or deleted is unsettled. Commenters on the first issue argued for keeping it with a status field. The book pulls the other way for memories about people, which are personal data: “be able to delete on request”. My policy block keeps superseded records out of the read path and leaves deletion as a line to fill in.

Provenance makes cleanup possible: when a wrong fact surfaces, the source field tells you what else that source wrote. I found no published measurement of a good expiry period, so every number in the policy block is a value to set and test.

How do you know memory helped?

You know memory helped when a task set scores better with it than without it, and when four specific failures stay absent. Chapter 9 separates the two claims that get confused: “we added memory” and “the agent is coherent across sessions” are, in its words, “different claims, and only the second one matters”. Du’s survey makes a related complaint: “Nobody evaluates forgetting well.”

Public benchmarks measure recall more than writing, as their own abstracts show.

Benchmark What it measures What it leaves out
LoCoMo (2024) Question answering, event summarization and dialogue generation over long generated conversations Agents that act; whether a write was correct
LongMemEval (2024) Five abilities of chat assistants, including knowledge updates and abstention, over 500 questions Decisions during a task; poisoning; upkeep cost
MemoryAgentBench (2025) “accurate retrieval, test-time learning, long-range understanding, and selective forgetting” Multi-session tasks with consequences
MemoryArena (2026) Memory used across multi-session tasks with interdependent subtasks I read the abstract only and quote no results

One number is worth its setup. In the LongMemEval paper (arXiv:2410.10813 v2), long-context models of 2024 read a full chat history of about 115,000 tokens. Against a setting that supplied only the sessions holding the evidence, they “showed a 30% to 60% performance decline”. The result supports selecting over pasting, for question answering in chat and for models of that year.

The four tests below are mine and cover what those benchmarks leave to you. Run them as a regression gate on every change to the write policy.

  • Forced reset. Kill the window mid-task and restart from notes. Pass: the task succeeds at the rate it did without the reset.
  • Changed-fact pair. Store a fact, then a confirmed correction, then ask. Pass: only the correction is used.
  • Cross-owner leak. Write a marker for one owner and query as another. Pass: the marker never appears.
  • Planted-instruction write. Put a harmless marker instruction in an untrusted test document. Pass: no durable record contains it, and no rule changed.

Worked example: a support agent that remembers users

A customer-support agent that remembers users across weeks and reads inbound email lands on the last row of the decision table: store plus retrieval, with the full write policy.

Taxonomy. The agent needs all four kinds. Working state is each conversation. Episodic memory is the log of past tickets. Semantic memory holds per-user facts such as a preferred contact channel. Procedural memory is the support playbook, which only a person edits.

Decision table. The longest thing to remember is a user across sessions, for many users, and inbound email is content anyone can write. The row gives a store, an owner filter and no direct writes from outside content.

Write policy. Email is listed as untrusted and is read by a worker with no memory tool. Suppose an email says the customer now prefers text messages. The write step stores “claimed by email of [date]: prefers SMS”, scoped to that user, with the status claimed. The older “prefers email” record stays active until the customer confirms the change in a session, and only then becomes superseded. Until then a read returns “prefers email”, with the claim labeled beside it. A sentence in the same email that addresses the agent is returned as evidence and stored nowhere.

The three artifacts agree on one agent memory architecture, and the tests follow from it: the changed-fact pair uses the contact channel, and the leak check uses two customers. Two more land elsewhere. A one-session coding agent on a trusted repository lands on the first row: a notes file, no store, and a person merging any proposed rule. A multi-day research agent on the open web lands on the fourth: a chain of short runs whose notes hold sourced claims and expire with the task.

Where does this advice stop?

This advice stops where my evidence does: the designs and rules come from one book chapter and a handful of papers, and the two tables and the policy are untested constructions of mine. Treat the decision table as a starting position that your evals confirm or overturn.

I read full text for four papers (Sumers, Du, Wu, Dong) and only abstracts for the rest. No source I found measures a safe size for the always-loaded tier, or how often a checked summary still carries a misleading claim.

The quarantine extension inherits the book’s caveat. A worker deceived by what it read returns a deceived summary, and Chapter 9 calls the result “safer, never safe”. Multi-agent systems that share one memory add problems this post does not cover, beginning with who may write when an orchestrator-worker pattern gives several workers one store.

Some agents need none of this. A single call, or a short session with no user to remember, has working memory and at most an instructions file that a person edits, which is the first two rows of the table. “No agent-written memory, by decision” is a complete design note.

The write is the design

An agent memory architecture is decided at the write. Reading is the same retrieval everywhere, while the kinds, the designs and the policy differ in who may write, what is checked first, and how a wrong entry gets found. Rank your writes by the damage they can do, give the riskiest one a human reviewer, and test the four failures before you claim the agent remembers.

This post is one step in the write, select, compress and isolate sequence of context engineering, where write hands off to memory. The glossary is free to read, and its entries on memory and quarantine carry the book’s one-paragraph versions. Chapter 9, “Memory: Working State Across Long Runs” is in the full book, with the taxonomy, the tiers, the long-run protocol and the two-door rule at full length. The context engineering guide maps this cluster, the explainer on context rot and the four operations shows why windows fill, or you can see the formats.

Questions readers ask

What are the types of memory in an AI agent?
Four, in the taxonomy Chapter 9 of AI Agents, Engineered uses. Short-term or working memory is the context window itself. Long-term memory splits into episodic (what happened), semantic (what is known) and procedural (how things are done here). Sumers and colleagues formalized the four names for language agents in 2023, with a broader definition of working memory.
Is agent memory the same as RAG?
The read path is the same machinery. Chapter 9 calls a memory store one more corpus and calls reading from it retrieval-augmented generation over your own history. Memory adds two steps that a document corpus rarely needs: a write step, where the agent or a background process decides what to keep, and a manage step, where records are merged, superseded and expired.
Does a bigger context window replace agent memory?
A bigger window delays the question for one session and does nothing for facts that must outlive the session. On the LongMemEval benchmark (Wu and colleagues, 2024), long-context models of that year lost 30% to 60% of their performance when they read a full chat history of about 115,000 tokens, compared with reading only the sessions that held the evidence.
Should agent memory live in files or in a vector store?
Start with a file unless the agent serves many users. A file can be read, diffed and reviewed, and Chapter 9 calls it the baseline. Move to a store when the notes outgrow what a session can load or when memories must be filtered by owner. One 2026 survey gives the opposite default and recommends starting with a retrieval store.
What is memory poisoning?
Memory poisoning is content written to a durable store, by accident or by an attacker, that later sessions treat as true or as an instruction. The OWASP GenAI Security Project lists it as ASI06, Memory & Context Poisoning, in its 2025 Top 10 for Agentic Applications. The defense sits on the write path: restrict who may write, check what is written, and record where each record came from.

Sources

  1. Sumers, Yao, Narasimhan, Griffiths (2023). Cognitive Architectures for Language Agents (arXiv:2309.02427 v3; TMLR 2024)
  2. Pengfei Du (2026). Memory for Autonomous LLM Agents: Mechanisms, Evaluation, and Emerging Frontiers (arXiv:2603.07670 v1)
  3. Packer, Wooders, Lin, Fang, Patil, Stoica, Gonzalez (2023). MemGPT: Towards LLMs as Operating Systems (arXiv:2310.08560 v2)
  4. Dong, Xu, He, Li, Tang, Liu, Liu, Xiang (2025). Memory Injection Attacks on LLM Agents via Query-Only Interaction (arXiv:2503.03704 v5)
  5. Chen, Xiang, Xiao, Song, Li (2024). AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases (arXiv:2407.12784 v1)
  6. Johann Rehberger (2024). Spyware Injection Into Your ChatGPT's Long-Term Memory (SpAIware)
  7. OWASP GenAI Security Project (2025). OWASP Top 10 for Agentic Applications (announcement, 2025-12-09)
  8. Wu, Wang, Yu, Zhang, Chang, Yu (2024). LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory (arXiv:2410.10813 v2; ICLR 2025)
  9. Maharana, Lee, Tulyakov, Bansal, Barbieri, Fang (2024). Evaluating Very Long-Term Conversational Memory of LLM Agents (LoCoMo; arXiv:2402.17753 v1)
  10. Hu, Wang, McAuley (2025). Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions (MemoryAgentBench; arXiv:2507.05257 v4)
  11. He, Wang, Zhi, Hu, Chen, Yin, Chen, Wu, Ouyang, Wang, Pei, McAuley, Choi, Pentland (2026). MemoryArena: Benchmarking Agent Memory in Interdependent Multi-Session Agentic Tasks (arXiv:2602.16313 v2)
  12. Chhikara, Khant, Aryan, Singh, Yadav (2025). Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory (arXiv:2504.19413 v1)
  13. Rasmussen, Paliychuk, Beauvais, Ryan, Chalef (2025). Zep: A Temporal Knowledge Graph Architecture for Agent Memory (arXiv:2501.13956 v1)
  14. Anthropic (2026). How Claude remembers your project (Claude Code documentation, read 2026-10-07)
  15. AGENTS.md contributors (2026). AGENTS.md (project site, read 2026-10-07)
  16. utkucanaytac (GitHub) (2026). ADD-only memory extraction can create conflicting memories (GitHub issue, open)
  17. empyreaaron (GitHub) (2026). Search ranks an outdated memory above its update (GitHub issue, open)
  18. silentsvn (Hacker News) (2026). Hacker News comment on contradictory facts recalled with equal confidence
  19. harperlabs (Hacker News) (2026). Hacker News comment on file-based memory going stale after three weeks
  20. kaydub (Hacker News) (2026). Hacker News comment on agents that amend and never delete
  21. JohnBooty (Hacker News) (2026). Hacker News comment on rediscovery in every session