An agent is only as good as what it can find. You can put every answer your company owns into one place, but if the agent has to read the whole thing to answer one question, you've built a haystack, not a brain.
So we built for retrieval first. Every document an agent writes into Tatara carries structure a machine can act on: a type, tags, a one-line description, and a place in a map of the whole brain. When an agent needs your refund policy, it goes straight to the one page that matters instead of scanning a thousand.
Why can't an agent just search everything?
An agent can search everything, but search returns candidates, and reading candidates is where the cost is. A full-text query for "refund" across a real brain comes back with the policy, four meeting notes that mention refunds in passing, a pricing page and a support macro. Now the agent has to open some of them to find out which one it wanted.
Each of those reads costs tokens and latency, and every irrelevant document it reads is a chance to ground its answer on the wrong page. An agent that reads ten documents to answer one question is an agent people stop using, because it's slow and occasionally confidently wrong. Structure is there to get that number down to one.
How does structure make retrieval work?
Structure tells an agent what a document is before the agent reads a word of it. Search alone gets you close, and structure gets you the right answer. A tagged, typed, described document can be found in a single hop instead of a trawl through everything that mentions "refund."
You can't add findability later. It decides whether your agents use a knowledge base or quietly route around it.
What signals does every document carry?
Documents an agent writes into Tatara carry four signals, and people can set the same four on their own. Each one narrows the search in a different way.
- Type. Free-form and declared on every agent write:
policy,playbook,meeting-note,icp. An agent looking for a rule can filter to the types that hold rules and ignore the fifty meeting notes that merely discuss them. - Tags. One to five per document, from a three-tier registry: tier 1 is the area, tier 2 the topic, tier 3 the goal. Tags cut across folders, which is what you need when the answer lives in Sales and the question came from Support.
- Description. One sentence saying what the document is. It appears in listings, so an agent can rule a document in or out without fetching its body.
- Folder path. Where the document sits. It scopes a search to one part of the brain and gives a person somewhere sensible to browse.
None of these are inferred after the fact. They're required at the moment an agent writes, which is the only time anyone reliably knows the answers.
What does the map look like to an agent?
To an agent, the map is a single tool call. get_taxonomy returns the brain's current shape: the real folders with their slugs, names, descriptions and full paths; the document types actually in use; the tag registry with its three tiers; and the registry of custom fields. Agents call it once at the start of a session and cache the result.
With that call, an agent knows your structure instead of guessing at it. Without a map, an agent writes into a folder it invented and tags with vocabulary nobody else uses, and the brain degrades one well-meaning write at a time. With a map, it files things where your team files things, because it can see where that is.
The map reports what exists rather than what a schema wishes existed. A new brain with no folders and no tags returns empty lists, and empty is the correct answer.
Wouldn't better search solve this instead?
Better search matching solves a different half of the problem. Tatara's search is Postgres full-text search over the document body: you ask for "refund window" and it ranks the documents that talk about refund windows. Swap in any smarter matching you like and it will rank the same set more cleverly. It still can't tell you which of those documents your company actually stands behind.
Similarity answers "what looks like your query." Structure answers "what is this, who owns it, is it current, and is it the canonical one." A meeting note where someone floated changing the refund window is an excellent semantic match for a question about the refund window, and a terrible answer to it. No amount of ranking fixes that, because the ranking has nothing to rank on. The note and the policy differ in what they are, not in their words.
Declared structure survives rewording, too. Retitle the policy and rewrite half its sentences, and it is still typed policy, still tagged billing, still described as the canonical one. A retrieval strategy built on those signals keeps working through edits that would move every similarity score in the corpus.
A worked lookup
Here is the actual sequence for "what's our refund window for annual plans?"
- Orient.
get_taxonomyreturns the folders and the tag registry. The agent sees there's apoliciesfolder and abillingtag. - Search, scoped.
search_documentswith the query plus a folder or tag filter. It comes back with ranked hits, each carrying highlighted snippets, an id, a path, a type and its tags. Filters compose with the query and apply before the limit, so narrowing really narrows instead of just re-sorting. - Triage without fetching. The hits and listing rows carry the one-line description: "The refund and cancellation policy for all paid plans, including annual." The agent knows which document it wants before reading any document.
- Fetch one.
get_documenton that id. The agent reads the current text and answers from it.
That's four steps and one document read. When the agent has no search terms at all ("what changed in billing this month?"), it uses list_documents with an updated-since timestamp instead. That returns lightweight rows and never content, so browsing the brain stays cheap even when it's large.
Without the description line, step 3 disappears and the agent fetches three or four full documents to work out which one it wanted. Everything still works. It just costs several times as much, and it gets slower as you write more. Findability is what keeps a brain from getting worse as it grows.
Written for people, mapped for machines
None of this shows up as clutter for your team. People see clean documents and a sensible shelf. Agents see the same documents plus the map underneath: folders, tags, and links that say how everything connects. It's one artifact with two ways in.
Rendering follows the same rule. A document can declare a render format and appear to your team as a data table, a kanban board, a task tracker or a mind map, while the stored content stays Markdown that an agent parses directly. The board your team drags cards around on and the text the agent reads are the same file. Nothing gets converted between the two, and there is nothing to keep in sync.
How to make a brain findable
Most of the work is habit, not configuration.
Write descriptions that discriminate. "Pricing" is useless; "how we set and approve discounts on annual contracts" tells an agent whether to open it. Assume the description is the only thing anyone reads.
Keep folders shallow and few. Deep trees express relationships that tags express better, and every extra level is another chance to file something where nobody looks for it.
Register your tag vocabulary rather than letting it accrete. Three tiers is enough: the area, the topic, the goal. A registry of twenty tags everyone uses beats two hundred that each appear twice.
Give a document one job. A single page covering pricing, packaging and discount approvals is one hit for three different questions, and the wrong answer for at least two of them.
The payoff is unglamorous: agents find the right document on the first try, because the brain was built to be searched as well as stored.
