First Principles

Company brain glossary: the vocabulary of the context layer

Plain definitions of the terms behind company brains and context layers: governed write, provenance, MCP, RAG, grounding and thirteen more.

AngusFounder of Tatara · Aug 29, 2026 · 11 min read

This glossary covers the vocabulary that has grown up around AI agents and company knowledge: eighteen terms, in alphabetical order. Each is defined the way the field uses it, not the way a vendor would prefer. Where a term has an accepted meaning, that's the meaning given here, and where it has no settled meaning yet, the entry says so. Where Tatara means something specific by a word, the general definition comes first and the implementation detail after it.

Agent memory

Agent memory is the mechanism an AI assistant uses to keep information about a person or a task between separate conversations, so a later session can use something an earlier one established. The term has no settled technical definition. It gets applied to at least three different things: a per-user note store inside one product, a running summary of past sessions, and a vector index over earlier transcripts. In all three, the store belongs to the assistant rather than the organisation, so it's scoped to one product and usually to one person. See also retrieval, which comes at an overlapping problem from the other direction.

Audit trail

An audit trail is an append-only record of the changes made to a system: what changed, who or what changed it, and when. It answers a different question from provenance. Provenance says where a document came from, and the audit trail says what has happened to it since. That second question is the one you need answered when somebody disputes a fact and you have to reconstruct how it got there. In Tatara, every document and folder mutation writes exactly one audit row inside the same database transaction as the change, so a mutation cannot commit without its record.

Company brain

A company brain is a single, maintained store of what an organisation knows about itself, kept in a form that people can read and AI agents can retrieve as working context. It holds the operational knowledge a business runs on and rarely writes down properly: who the ideal customer is, how pricing works, how a deal moves, what was decided in March and why. It differs from a folder of documents in one way that matters: something downstream depends on it staying true, because an agent reads it and acts on it. In Tatara, "brain" is the name for that store. A workspace holds exactly one, and every document, folder and tag belongs to it.

Context engineering

Context engineering is the practice of deciding what information a model sees for a given task, in what form, and in what order. It replaced "prompt engineering" as the usual framing during 2025, once it was clear that most failures in agent systems came from what was in the context window, not from how the instruction was worded. In practice it means choosing which documents to retrieve, compressing them to fit, ordering them so the load-bearing material isn't buried at the back, and deciding what to leave out. A context layer is the infrastructure that practice reads from.

Context layer

A context layer is a shared store of an organisation's knowledge that sits between its information and the AI systems reading it, so agents draw from one maintained source instead of from whatever was pasted into each conversation. The term isn't settled, and vendors use it for quite different things: a vector database, a documentation product with an API, a per-assistant memory service. What they have in common is a single addressable place that agents read from. Implementations differ on whether writes into it are constrained and attributed, which is what governed write covers.

Context window

A context window is the amount of text, measured in tokens, that a language model can consider in a single exchange: the instruction, any retrieved material, the conversation so far, and the model's own output. It's working memory for one conversation, not storage. Nothing put into it survives the conversation closing, and no other tool or person can see what's in it. A larger window gives that working memory more room without making it last, which is why "the model forgot" usually describes a missing retrieval step rather than a limit of the model.

Frontmatter

Frontmatter is a block of structured metadata at the top of a text file, conventionally written in YAML and fenced by lines of three hyphens, that describes the document below it. It came out of static site generators and became the default way to attach machine-readable fields to a human-readable document without adding a database. A typical block carries a title, a type, a one-line description, tags and dates. Frontmatter is what makes a plain Markdown file addressable by a machine. The prose is for the person, and the fields are what an agent reads first to decide whether this is the document it wants.

Governed write

A governed write is a write into a shared knowledge store that is validated against required structure and attributed to its author at the moment it lands, rather than accepted freely and reviewed later. This matters now that people aren't the only writers. A person writes slowly, rarely and with their name attached; a process can produce two hundred documents overnight under no name at all. In Tatara, an agent creating a document must supply a type, a folder, a title, a body, a one-sentence description and one to five tags. A write missing any of them is rejected with an error that explains how to fix it. Every save is stamped with provenance.

Grounding

Grounding is constraining a model's answer to material retrieved from a specified source, so the response reflects that source rather than the model's own parametric knowledge. An ungrounded model answering about your company works from general patterns in its training data. That produces fluent, confident, generic answers, and where the pattern doesn't fit, fabricated ones: the failure usually called hallucination. Grounding doesn't make a model more capable. It changes what the model is answering from, so the source sets the ceiling on quality. An agent grounded in a stale document is exactly as wrong as the document, and a good deal more persuasive about it.

Knowledge base

A knowledge base is an organised collection of information about a subject, maintained so it can be looked up rather than rediscovered. In software the phrase almost always means one of two things: a customer-facing help centre or an internal wiki. Both are written for a human arriving with a question, which shapes them into long explanatory pages rather than retrievable units, and neither usually attaches an owner to any given page. A company brain differs on both counts: it's organised around what a machine will fetch, and around who is accountable for it being true.

Markdown

Markdown is a plain-text formatting syntax, created by John Gruber in 2004, that uses ordinary punctuation for structure: hashes for headings, asterisks for emphasis, hyphens for lists. Two things make it matter here, and neither has anything to do with how it looks. Language models are trained on it and write it more reliably than any other format, so knowledge stored as Markdown needs no translation step before an agent can use it. And it stays legible without the software that produced it: a Markdown file opens in any text editor, diffs cleanly in version control, and can be taken elsewhere intact.

Model Context Protocol (MCP)

The Model Context Protocol (MCP) is an open protocol, published by Anthropic in November 2024, that gives AI applications a uniform way to connect to external tools and data sources. A server exposes a set of tools with self-describing schemas, and a client such as Claude, ChatGPT, Gemini or a coding editor discovers those tools and calls them mid-conversation. It matters because it turned a per-product integration problem into a single interface: a system that speaks MCP is reachable from every MCP-capable client without bespoke work at either end. Tatara runs an MCP server with 23 tools, covering search, reading, writing, filing, taxonomy curation and restoring deleted documents.

Open Knowledge Format (OKF)

The Open Knowledge Format is an open specification published by Google Cloud for representing organisational knowledge as plain Markdown files with YAML frontmatter, intended as a portable substrate for the context AI agents read. It formalises a shape that people doing this by hand had already settled on: one file per concept, structured fields at the top, links between files. The specification is young and still at a 0.x version, so it's better read as a sign of the direction the industry is standardising in than as a compliance target. Tatara's documents converge on the same substrate: Markdown bodies, structured frontmatter and no proprietary encoding.

Provenance

Provenance is the record of where a piece of information came from: who or what created it, under which credential, and when, written by the system rather than supplied by the writer. The last part is what matters. If a client can send its own author field and have it stored, you don't have provenance, only a self-declaration that happens to be saved. Once agents write at volume, the property worth having is that machine-authored content can always be told apart from human-authored content. In Tatara the stamp is derived from the token that authenticated the call, and it cannot be supplied, overridden or faked by the writer.

Retrieval

Retrieval is fetching a specific document or passage at the moment it's needed and placing it in the model's context, so the answer comes from that text rather than from what the model already believes. It handles the durable half of the memory problem. What persists is the document, not the conversation, so every agent and every person reads the same current version. That's the difference from agent memory, which persists inside one product for one user. How well retrieval works depends on what an agent can see before it opens anything, which is usually a title, a type, a description and tags.

Retrieval-augmented generation (RAG)

Retrieval-augmented generation (RAG) is a technique, introduced in a 2020 paper by Lewis et al., in which a system searches a corpus for passages relevant to a query and places them in the model's context before it generates an answer. In its common form the corpus is split into chunks, each chunk is embedded as a vector, and the query retrieves the nearest ones. RAG is a way of finding material, so it assumes there is material worth finding. Point it at a corpus of stale, unowned and contradictory documents and it returns those documents quickly, in chunks stripped of any signal about which one was current.

Source of truth

A source of truth is the one place a fact is authoritatively recorded, and every other copy of that fact defers to it. The phrase comes from data management, where "single source of truth" describes a system designed so a value is stored once and referenced everywhere else instead of being duplicated. Applied to company knowledge, it's more a governance claim than a technical one. It means a named person is accountable for keeping this document true, and that when it changes, nothing downstream has to be updated in turn. A store copied into three other systems has three sources of truth and no way to tell which is current.

Taxonomy

A taxonomy is the controlled vocabulary a knowledge store uses to classify its documents, so the same idea is filed under the same word every time. Without one, a store collects four spellings of the same concept, and the answer to "everything about pricing" quietly leaves out half of it. The usual tension is between letting writers coin terms freely, which produces sprawl, and requiring approval first, which stalls them. Tatara's tag registry sorts tags into three independent tiers: Area (who owns it), Topic (what it's about) and Goal (what it moves), with up to five tags per document. Agents may coin unregistered tags, which land in a review queue for a person to file or merge.

THE INVITATION

Give your agents a brain.

Start free and build the one source of truth every AI you use can read from.

Start freeVisit the homepage →
No credit card · by invitation
Angus

Founder of Tatara. Writes about company knowledge, AI agents that do real work, and building a brain a whole team can trust.

KEEP READING

More from the Ledger