Context Engineering for AI Agents: Memory, Retrieval, and Working Context
A practical architecture guide to context engineering for AI agents: what belongs in working context, what should be retrieved, what should persist as memory, and when to trim or compact state.
On this page
The problem with agent context is usually not that the model has too little information.
It is that the system has not decided which information deserves to be active right now.
An agent can have access to a large context window, a vector database, a memory store, dozens of tools, a long conversation history, and live application state—and still perform poorly because too much irrelevant information reaches the model at the wrong time.
That is the core problem context engineering tries to solve.
Prompt engineering asks how instructions should be written. Context engineering asks a broader systems question:
What information should the model see at this step, where should that information come from, how long should it remain available, and what should be discarded?
For agents, this distinction matters because context is not static. Every tool call, observation, user message, retrieved document, plan, error, and intermediate result can become candidate context for the next turn.
The engineering task is therefore not to maximize context. It is to curate working state.
A useful mental model: four different things
The terms context, memory, retrieval, and history are often used interchangeably. They should not be.
A practical architecture separates at least four layers.
1. Working context
Working context is what the model can directly attend to during the current inference step.
It may include:
- system instructions;
- the current user request;
- recent conversation turns;
- tool definitions;
- selected tool results;
- retrieved documents;
- a current plan or scratch state;
- environment observations.
This is the model’s active workspace.
2. Retrieval
Retrieval is the mechanism that decides which external information should be brought into working context.
That information may come from:
- documentation;
- databases;
- code repositories;
- knowledge bases;
- prior conversations;
- search indexes;
- files;
- external APIs.
Retrieval does not automatically imply memory. A system can retrieve a product manual without remembering anything about the user.
3. Persistent memory
Memory is information intentionally preserved across turns, sessions, or tasks because future behavior may benefit from it.
Examples:
- a correction the user made;
- a stable project constraint;
- a decision already approved;
- a known failure mode;
- a preference that should affect later behavior;
- a compact handoff from a previous long-running session.
Memory is a write policy as much as a read policy. A useful memory system must decide what is worth saving and when it is worth loading again.
4. External state
Some information should not be represented as model memory at all.
The authoritative state may already live in:
- a database;
- Git;
- an issue tracker;
- a calendar;
- an account system;
- a payment ledger;
- a running application;
- a filesystem.
If the agent can inspect the source of truth when needed, copying that state into long-term memory can make the system less reliable rather than more reliable.
The first design rule is therefore simple:
Do not use memory to duplicate a source of truth the agent can query directly.
Larger context windows do not remove the architecture problem
It is tempting to assume that context engineering becomes less important as context windows grow.
In practice, larger windows reduce one constraint while leaving several others intact.
Anthropic describes context as a finite attention resource and argues for the smallest set of high-signal tokens that improves the probability of the desired behavior. Its context-engineering guidance also highlights a familiar problem: long contexts can accumulate irrelevant or stale information even when they technically fit inside the model’s window.
OpenAI’s session-memory guidance makes a similar point from an implementation perspective. Long-running agents can carry too much history forward, introducing distraction, latency, and cost. The recommended tools include trimming older turns and compressing history into summaries rather than treating every previous token as equally valuable.
The important distinction is:
Capacity answers how much can fit. Context engineering answers what should be present.
Those are different questions.
Context should be assembled, not accumulated
A naive agent loop often works like this:
- append the latest message;
- append the tool call;
- append the complete tool result;
- append the next message;
- repeat indefinitely.
This treats the transcript as the agent’s state model.
That works for short tasks. It becomes fragile for long-running work.
A better architecture assembles context for each step from several controlled sources.
For example:
- stable system instructions;
- the current objective;
- a compact representation of progress;
- only the recent turns that still matter;
- tool definitions relevant to the current phase;
- retrieved evidence relevant to the current question;
- live state from authoritative systems when needed.
The transcript can still be stored for observability and auditability without forcing the model to reread all of it on every turn.
Preload only what is predictably useful
Some context should be available from the beginning.
Good preload candidates include:
- task-defining instructions;
- safety and permission boundaries;
- output requirements;
- a small number of stable project constraints;
- tool contracts the model must understand immediately.
Bad preload candidates include:
- entire documentation sets;
- every previous task outcome;
- every tool the system supports;
- large logs;
- full repositories;
- all historical user interactions.
The more uncertain the relevance of a piece of information, the stronger the case for retrieving it later rather than preloading it.
This is why Anthropic’s current context-engineering guidance emphasizes just-in-time retrieval: keep lightweight references available and let the agent load deeper information when the task actually requires it.
Retrieval should answer a concrete question
RAG is often treated as a generic solution for context.
But retrieval quality depends on what the system is trying to resolve.
Before retrieving, ask:
- What uncertainty is the agent trying to reduce?
- What source would be authoritative for that uncertainty?
- How much evidence is enough?
- How fresh must the evidence be?
- What permissions apply?
Suppose an agent is debugging a production incident.
Useful retrieval might include:
- the current error trace;
- the relevant service configuration;
- recent changes to the affected component;
- the runbook for that service.
Retrieving twenty generic architecture documents because they are semantically similar may make the model less focused.
Anthropic’s Contextual Retrieval work addresses one part of this problem: document chunks can lose meaning when separated from their surrounding document, so adding compact document-level context before indexing can improve later retrieval.
The larger principle is more important than the specific technique:
retrieval should preserve enough meaning for the returned evidence to be useful once it reaches working context.
Memory should preserve hard-to-reconstruct information
A persistent memory system is most valuable when the information is important and expensive to infer again.
OpenAI’s internal data-agent architecture provides a useful example. The agent stores non-obvious corrections, filters, and constraints that are important for querying internal data correctly. It does not rely on memory as a copy of the whole data warehouse. Live data and metadata remain queryable from their authoritative systems.
That distinction is a good design rule.
Good memory candidates are often:
- user corrections;
- project-specific terminology;
- stable constraints;
- decisions and rationale;
- failure lessons;
- durable preferences;
- compact task handoffs.
Weak memory candidates are often:
- temporary API responses;
- current inventory counts;
- live account balances;
- raw logs;
- rapidly changing documentation;
- information that can be cheaply and authoritatively queried again.
The cost of bad memory is not only storage.
Stale memory can actively push the agent toward the wrong action.
Memory needs write rules
The most dangerous memory architecture is “save everything that seems useful.”
That creates a second uncontrolled history.
A stronger system defines explicit write criteria.
Before persisting something, ask:
- Is it likely to matter again?
- Will it remain true long enough to be useful?
- Would reconstructing it later be expensive or unreliable?
- Does saving it introduce privacy or permission risk?
- Can the original source of truth be queried instead?
You can think of memory as a cache with semantic consequences.
Caching the wrong thing is not neutral.
Long-running agents need compaction
Eventually even a well-curated active context can grow too large.
Compaction is the process of replacing a large amount of prior context with a smaller representation of the information that should survive.
A useful compacted state might preserve:
- the objective;
- completed work;
- decisions already made;
- unresolved problems;
- current hypotheses;
- files or resources that matter;
- validation already performed;
- explicit next steps.
It should discard things such as:
- redundant tool output;
- superseded hypotheses;
- repeated acknowledgements;
- verbose intermediate reasoning that no longer affects the task.
Anthropic describes compaction and structured note-taking as key techniques for long-horizon agents. Its more recent managed-agent work goes one step further and separates the persistent session from the model’s current context window, because not every historical token should have to remain active for the work to continue.
That is a useful architectural boundary:
the session is the full durable record; the context window is the temporary working set.
Trimming and compaction are different
These techniques are related but not identical.
Trimming removes material according to a rule.
Examples:
- keep only the last N turns;
- drop old tool results;
- remove previous reasoning blocks;
- retain only messages after the most recent checkpoint.
Compaction transforms material into a smaller representation.
Examples:
- summarize completed work;
- convert a long interaction into a structured handoff;
- persist decisions into a task state object;
- reduce multiple tool traces into verified conclusions.
Trimming is cheaper and safer when old content is clearly disposable.
Compaction is better when information must survive but raw history is too expensive or distracting to preserve.
A robust system usually uses both.
Tool outputs deserve aggressive context policy
Tool results are one of the fastest ways to pollute an agent’s context.
A single browser page, SQL result, terminal log, or API response can consume thousands of tokens.
Instead of carrying raw results indefinitely, extract what the next step actually needs.
For example:
- browser page → relevant facts + source URL;
- test output → failing tests + important stack traces;
- database query → result rows needed for the decision;
- repository search → relevant paths + excerpts;
- API response → state transition + identifiers.
Keep the raw artifact somewhere retrievable when auditability matters.
Do not force it to remain in working context just because it once existed.
Freshness and permissions are part of context engineering
Correct context is not only relevant. It also has to be current and authorized.
A memory written six months ago may conflict with today’s database state.
A retrieved document may be relevant but not visible to the current user.
A tool result may expose more information than the next model turn needs.
Every context source therefore needs metadata such as:
- source;
- timestamp or version;
- ownership;
- permission scope;
- confidence or verification state;
- expiration policy when appropriate.
This becomes especially important when agents combine organizational knowledge with user-specific information.
Retrieval without access control is a security bug.
Memory without freshness policy is a reliability bug.
A practical context architecture
A production agent can use a layered structure like this:
| Information | Default location | Load policy |
|---|---|---|
| System rules | Working context | Always |
| Current user request | Working context | Always |
| Recent interaction | Working context | Bounded window |
| Current plan/progress | Structured working state | Always while task is active |
| Documentation | Retrieval index | On demand |
| Live application state | Source-of-truth system | Query on demand |
| Durable user/project constraints | Persistent memory | Retrieve when relevant |
| Raw historical tool results | Artifact/log store | Only when needed |
| Long-session history | Durable session record | Compact/trim before model use |
The exact components will vary, but the separation matters.
When everything becomes “memory,” nothing has a clear lifecycle.
Common failure modes
Context stuffing
The system loads everything that might be relevant.
Result: high token cost, distraction, and weaker attention to the information that actually matters.
Stale memory
Old information is treated as authoritative after the real system has changed.
Result: confident decisions based on obsolete state.
Retrieval noise
Semantically related documents are retrieved even though they do not answer the current question.
Result: plausible but unfocused reasoning.
Summary drift
Repeated compression slowly removes details or changes their meaning.
Result: the agent remains coherent but becomes coherent around an inaccurate representation of the task.
Tool-result accumulation
Every large observation stays in the conversation forever.
Result: useful context is buried under operational exhaust.
Memory overreach
The system persists information that should have remained temporary or private.
Result: privacy, permission, and behavioral problems across future sessions.
Evaluate context policy as part of the agent
Context engineering should be evaluated, not treated as invisible plumbing.
Useful experiments include:
- full history vs. trimmed history;
- preloaded documents vs. just-in-time retrieval;
- raw tool results vs. extracted observations;
- no memory vs. structured memory;
- different compaction strategies;
- fresh vs. intentionally stale memory injections;
- retrieval with and without metadata filtering.
Measure downstream outcomes:
- task success;
- latency;
- token consumption;
- tool-call count;
- hallucination or grounding failures;
- recovery after long sessions;
- policy violations;
- unnecessary retrievals.
The goal is not to minimize tokens for its own sake.
The goal is to maximize useful evidence per token of active context.
The design rule
Treat context as a working set, not an archive.
Keep stable instructions small. Retrieve large knowledge just in time. Persist only information that is costly to reconstruct and likely to matter again. Query live systems instead of memorizing volatile state. Trim disposable history. Compact the information that must survive.
The strongest agent architecture is not the one that remembers everything.
It is the one that can reliably decide what to remember, what to retrieve, what to verify again, and what to forget.