"Agent memory" and "context engineering" get swapped for each other constantly, often inside the same sentence. They describe different problems. Mixing them up is how teams end up building a memory layer to fix a retrieval problem, or buying a bigger context window to fix a staleness problem.
The distinction fits in one line: agent memory is what a system keeps between calls; context engineering decides what goes into the next one.
Memory is a storage problem and context engineering is a selection problem. Get the storage right and the selection wrong, and the agent still gets it wrong, because a fact that never reaches the window may as well not exist.
What is agent memory?
Agent memory is the machinery that lets an agent carry something forward from one interaction to the next. Two quite different things share the name.
Short-term or working memory is the current session: the conversation so far, held in the context window. Nothing persists it, and when the session ends it is gone. A surprising number of features labelled "memory" are only this, plus a summarisation step when the window starts to fill.
Long-term memory is state written somewhere durable (a database row, a vector index, a file) and read back on a later call. This is the harder engineering problem, and it is what most people mean by the term.
Long-term memory is usually split further, using labels borrowed from cognitive psychology:
- Episodic memory records specific events. "On 14 August the user said the launch had slipped to Q4."
- Semantic memory holds the facts distilled out of those events, with the event itself discarded. "The launch is in Q4." It gains reusability and loses provenance.
- Procedural memory, where a system implements it, holds operating instructions the agent has learned and keeps updating about its own behaviour.
Mechanically, a write is usually an LLM call that extracts candidate facts from a transcript. The result is stored as an embedding in a vector store, or as structured rows keyed by user or thread. A read is a similarity search or a straight lookup.
These terms are not standardised. Two products can both advertise "long-term memory" and mean an append-only event log in one case and a continuously rewritten user profile in the other. They behave completely differently when a fact goes stale or gets contradicted, so ask which one you are getting.
What is context engineering?
Context engineering is the practice of deciding what a model sees on a given call, within a finite budget. It covers everything that occupies the context window, not only the prompt.
On a typical agent call, that window holds:
- The system prompt and operating instructions
- Tool definitions (names, descriptions and schemas), which cost tokens whether the tools get used or not
- Retrieved material: documents, search hits, database rows
- Tool results from earlier steps in the same run, which can end up crowding out everything else
- Prior conversation turns, in full or compacted
- The user's actual request
The work is retrieval, selection, ordering, compression and eviction. Ordering matters more than people expect. Models attend unevenly across a long context, and material buried in the middle is more likely to be missed than the same material near the start or the end.
The budget has changed the most. Context windows have gone from roughly 4,000 tokens in 2022 to hundreds of thousands today, and a million in some models. That moved the constraint from "will it fit" to "what deserves the space". More tokens do not buy more accuracy. Filling a window with plausibly related material lowers the signal ratio and makes the wrong answer easier to reach for.
MCP, the Model Context Protocol that Anthropic published in November 2024 and that has been widely adopted since, belongs squarely to this half. It standardises how tools and data sources reach a model, which makes it plumbing for context engineering rather than a memory system. Connecting an MCP server does not give an agent memory.
The term itself is recent. Context engineering came into common use through 2025, largely as people recognised that prompt engineering only ever described one part of the window.
How is agent memory different from context engineering?
Agent memory asks what should survive this session; context engineering asks what should be in the window on this call. The other differences follow from that.
| Agent memory | Context engineering | |
|---|---|---|
| The question it answers | What should persist beyond this session? | What should the model see on this call? |
| Unit of work | A fact or event, written once | A single model call, assembled every time |
| Time horizon | Across sessions, potentially years | The next few seconds |
| Typical mechanisms | Fact extraction, vector or key-value stores, profile records, summarisation | Retrieval, reranking, ordering, compaction, tool-result pruning |
| Characteristic failure | The agent forgets, or recalls something no longer true | The right information exists but never reaches the model |
| What fixes it | A durable store with an update path | A better selection and ordering policy |
The dependency is lopsided. Every memory system ends in a context-engineering decision, because a remembered fact does nothing until something puts it in the window. Plenty of context-engineering problems, though, have nothing to do with memory.
Where do agent memory and context engineering overlap?
Agent memory and context engineering meet at retrieval, and that is where most memory systems fail. A store is inert until something decides which of its contents matter for this call and where they go in the window.
The bottleneck is rarely recall. It is precision. Pulling thirty loosely relevant past turns is worse than pulling the three right ones, because the other twenty-seven crowd out the tool results and documents that would have produced a correct answer. Teams tend to measure whether the memory was found, and not whether it displaced something better.
The overlap runs the other way too. Context engineering with no durable store can only assemble from what is in front of it: this run's tool results and this conversation. That is fine for a one-shot task. It is useless for an assistant that is expected to know you moved the launch date last month.
Do you need both?
Whether you need both depends on one test: was the knowledge your agent is missing ever said to it? That settles most cases.
If the missing thing happened in a past conversation (a preference the user stated, a correction they made, a decision from last Tuesday's thread), you have a memory problem. Something needs to have written it down and be able to find it again.
If the missing thing lives in your company and was never said to the agent at all (your pricing tiers, your ICP, the escalation policy, last quarter's post-mortem), no memory system can recover it, because there is nothing to remember. That is a retrieval and context problem, and you solve it by giving the agent access to a current source.
Memory is for what the agent was told. Context engineering is for what your company knows. Most "the agent has no memory" complaints turn out to be the second problem wearing the first one's clothes.
In practice, a long-running personal copilot needs both, because it accumulates state about one person that exists nowhere else. A support agent answering from a policy base needs the context half done well and may need no long-term memory at all. Build the half you need first, and measure before you add the other.
What goes wrong when you confuse them?
Teams that confuse agent memory with context engineering build the wrong layer, and the symptom survives the fix. The failures tend to look like this:
- Memory that should have been a document. A team adds a memory layer so the agent stops getting the refund policy wrong. The agent never learns the policy, because nobody ever told it the policy in a conversation. What it needed was a readable source, not a store of past turns.
- A memory per tool. More and more assistants ship with memory of their own. Each one learns separately and none is authoritative, so contradictions between them never resolve, and the accumulated state does not follow you when the team switches tools.
- Remembering with no way to correct. An append-only store hands back a superseded fact with exactly the same confidence as the current one. A memory system without an update and delete path turns into a liability at the moment it starts being useful.
- Buying window instead of doing selection. A larger context window absorbs a selection problem for a while. Then it reintroduces it, at higher cost and latency, once the extra space fills up with near-misses.
- No provenance. Once material is in the window, something an agent inferred looks identical to something a person wrote and stands behind. If your agents can write to the same store they read from and nothing records which is which, the store degrades without anyone noticing, and you find out late.
Where does Tatara fit?
Tatara answers the context-engineering half, for the part of the problem that is organisational rather than personal. It is a governed document store that agents read from over MCP: what your company knows, in one place every assistant can reach.
Every agent write is governed. It needs a type, title, description and tags, so nothing lands as an unfiled note. Every write is also stamped with provenance, so an agent cannot pass as a person, and the record always says where each piece of material came from. Documents are portable Markdown, so the thing your agents depend on is something you can read, edit and walk away with.
It is not a memory system. It does not remember your conversations, and it will not tell you what you said last Tuesday. If that is your gap, you want a memory layer, and Tatara is the wrong tool.
Before you build either layer, check the simplest failure first. If the fact your agent got wrong was never written down anywhere, in a conversation or in a document, neither memory nor retrieval will save you. Write it down, then decide which layer needs to reach it.
