Product

What is the Open Knowledge Format? OKF explained

OKF is Google Cloud's open spec for storing what an organisation knows as Markdown files with YAML frontmatter. What it defines, and what it can't carry.

AngusFounder of Tatara · Aug 29, 2026 · 10 min read

The Open Knowledge Format (OKF) is an open specification, published by Google Cloud, for storing what an organisation knows as a directory of plain Markdown files: one file per concept, each carrying YAML frontmatter, cross-linked with ordinary Markdown links. It has no runtime, no SDK and no database. A bundle is a folder you could open in any text editor.

We pay attention to it at Tatara, a governed document store where a company keeps its business knowledge for its people and its AI agents to read, because our documents take the same shape and got there independently. Below is what the format specifies, what it leaves out, and where we sit relative to it.

What is the Open Knowledge Format?

The Open Knowledge Format (OKF) is a specification for writing organisational knowledge down as files instead of holding it inside a product. Google Cloud published v0.1 in June 2026 and v0.2 on 25 July 2026. The spec, sample bundles and reference implementations live in the GoogleCloudPlatform/knowledge-catalog repository.

The unit is a bundle: a directory of Markdown files, one file per concept. A concept is anything a reader needs to understand as a thing in its own right, such as a database table, a metric definition, a runbook or an API. A handful of rules do most of the work:

  • File path is identity. tables/orders.md is the Orders concept. There is no separate ID registry.
  • Every file is YAML frontmatter plus a Markdown body. Queryable facts go in the frontmatter; the explanation goes in the body.
  • type is the only required field. title, description, resource, tags and timestamp complete the v0.1 vocabulary, and all of them are optional.
  • Cross-references are plain Markdown links, so a folder becomes a graph with more connections than its own tree.
  • Two filenames are reserved: index.md for progressive disclosure, log.md for change history.

The spec deliberately says nothing about packaging. A bundle can be a tarball, a git repository or a folder mounted on a disk, and that gap is intentional, because a format with no runtime can't lock you in.

What problem does an open knowledge format solve?

An open knowledge format solves the problem of every AI tool wanting its own private copy of what your company knows. A model doesn't know your pricing, your ideal customer, or why you abandoned the old onboarding flow. Each tool that tries to fix that builds its own store, its own schema and its own idea of memory, and your knowledge ends up split across all of them.

A format moves the point of interoperability from the API to the file. An API is a dependency you have to keep, and a directory of Markdown isn't. Any tool that can read text can read the bundle, including tools that didn't exist when it was written.

Google's sample bundles show the intended shape. Google made them by pointing a reference enrichment agent at BigQuery public datasets (GA4 e-commerce, Stack Overflow, Bitcoin) and writing one Markdown file per table, enriched with schema, description and join paths. You can read the output without the agent that made it.

What changed in OKF v0.2?

OKF v0.2 added a vocabulary for deciding whether to trust a concept before reading it. v0.1 described what a concept is; v0.2 describes how much weight to put on it. The fields are organised around five questions:

QuestionFieldWhat it holds
What was this made from?sourcesAn array of {id, resource, title, author, usage_count, last_modified}; body claims cite them with Markdown footnotes
Who produced it, and who has confirmed it?generated, verifiedgenerated: {by, at} for the author; verified: [{by, at}] for independent confirmations. Author and confirmer are deliberately different fields
Is it still true?stale_afterOne absolute date, deliberately not a relative TTL, so staleness is a plain date comparison
Is this the current version?statusdraft → stable → deprecated. Absent means stable
Was this number produced the sanctioned way?Attested ComputationA runtime, typed parameters, an executor that returns a receipt, and an attester that verdicts the receipt. OKF records the computation and how to check it; it never executes anything

Three design choices in that release are worth pointing out, because you'll meet each of these arguments again elsewhere.

Decide before you read. Most interactions with a concept never reach the body, so the signals that decide relevance and trust have to be cheap and come first. That makes trust something you filter on before you spend tokens.

Signals, not scores. v0.2 records author, usage_count and last_modified, and declines to turn them into a credibility number. Scores are subjective, they don't carry over between organisations, and they go stale without anyone noticing. Consumers derive what they need.

Absence never rejects. Everything new is opt-in, type remains the only required field, and a v0.1 bundle is still a valid v0.2 bundle. Two things were superseded rather than broken: timestamp by generated.at, and the body's # Citations list by sources.

Why does Markdown with YAML frontmatter keep winning?

Markdown with YAML frontmatter keeps winning because it's the only common shape that serves a person and a machine from a single copy. The person reads the body and the machine reads the frontmatter. Nothing has to be converted or re-synced, so the human version and the machine version can't drift apart.

There's a less obvious reason, and it matters more than it should. Language models have read millions of tags: frontmatter blocks during training, so the convention is already in the weights. Calling the same field something else (topics, say) costs a sentence of explanation in every tool description forever and buys nothing. Writing knowledge in the shape models already recognise gets you accuracy for free.

The third reason is that there's nothing to implement. Every language already parses YAML and Markdown, and when you adopt a spec just by writing files, there's very little to abandon later.

What can a knowledge format not do?

A knowledge format can't make its own claims true. OKF frontmatter is self-declared, so anyone with write access to a file, including a careless or compromised agent, can type verified: [{by: human:ceo}] into it, and nothing in the format objects. v0.2 is candid about this: its derived trust tiers are advisory filters, never access control.

That's a real limit rather than a flaw. Whether a file is telling the truth is a question about the system that wrote it, and OKF has no provenance concept at all.

A format tells you what a file looks like. It cannot tell you whether to believe it.

Google made the second half of this argument itself. In an August 2026 post pairing OKF with its Knowledge Catalog product, it notes that a git repo per bundle "is not searchable alongside the data it describes, cannot be secured and governed using the same organizational identity and compliance policies," and does not scale. Google's answer is the format plus a context engine to hold it, which in its case needs Google Cloud, its IAM and a data engineering practice.

The format gets you portability and a shared vocabulary. Trust has to come from somewhere else.

See what governed knowledge looks like.Tatara is where your team writes down what your company knows, and every AI you use reads the same thing.
Start free →

What does portability actually mean in practice?

Portability is a property of the format your knowledge is stored in, not a feature a vendor bolts on later. If the substrate is Markdown and YAML, the worst case is a folder of files that any editor, agent or tool built in the next decade can read. If the substrate is a proprietary schema behind an API, portability is whatever that vendor's exporter chooses to emit this quarter.

People tend to blur two things here. A format can be open while the system holding it is closed, and a system can be generous about exports while still storing something only it understands. The useful question is what the bytes are while the knowledge is sitting there.

Tatara gives the same answer as OKF. Underneath, a brain is plain Markdown with YAML frontmatter, one file per concept, with links forming a graph. There's no proprietary schema, and nothing you couldn't open in an ordinary editor. That's a claim about the storage shape, and the storage shape is what decides whether lock-in is possible at all.

Where does Tatara sit relative to the Open Knowledge Format?

Tatara is convergent with the Open Knowledge Format, not conformant to it. We're careful about that distinction, and not embarrassed by it.

Tatara stores Markdown with YAML frontmatter, one file per concept, a path-based hierarchy, a type every agent write must declare, free-form vocabularies, and links parsed into a queryable graph. That shape was in the product before the spec existed. When Google published v0.1 we checked our architecture against it and found it already lined up, almost point for point. Two teams reaching the same answer without coordinating is better evidence that the shape is right than either of us designing it alone would be.

We ask for more than the spec does. OKF requires one field, and every agent write into a Tatara brain requires four: type, title, description and one to five tags, steered to that brain's own taxonomy. The spec deliberately defines the interoperability surface and leaves the content model to whoever produces the files, so a stricter house profile doesn't contradict it.

We also differ in places. We have no resource field: OKF's resource points at the thing a concept describes, while our source records who wrote the document, and those are different ideas. We removed status in July 2026 because nothing used it, three weeks before v0.2 defined draft, stable and deprecated. We don't implement verified, stale_after or Attested Computation.

So we don't describe Tatara as OKF-compatible, conformant or certified, and we won't until a round-trip-tested serializer exists to make the claim checkable. "Convergent" is the accurate word, and we'd rather use it.

The one place we're ahead of the spec is a field it doesn't have. A Tatara document's source is stamped from the authentication context and never typed into a file, so an agent provably cannot pass itself off as a person. Every agent write is governed on the way in, and every save is stamped with who or what made it. OKF records the trust signals. Making them true, and keeping them true, is the job of the system underneath.

Tatara is a standard format on the outside and a governed system on the inside. If OKF becomes the default shape for organisational knowledge, that costs us nothing, and it settles an argument we were already having.

THE INVITATION

Give your agents a brain.

Start free and build the one source of truth every AI you use can read from.

Start freeVisit the homepage →
No credit card · by invitation
Angus

Founder of Tatara. Writes about company knowledge, AI agents that do real work, and building a brain a whole team can trust.

KEEP READING

More from the Ledger