The pitch for an AI-maintained knowledge base is good enough that most people stop listening at the pitch. Point a model at your notes, let it write the summaries and keep the cross-references current, and the maintenance work humans always abandon gets done for roughly nothing. That part is real. What nobody puts in the pitch is the recurring bill: every question your agent answers from that corpus costs tokens to find the right page, and finding is a different problem from writing. We run one of these internally. In July we stopped guessing about the finding half and measured it.
Two terms, plainly. An LLM second brain is a folder of notes that a model writes and maintains for you, so the synthesis happens once instead of on every question. Index-first retrieval means the agent scores a one-line-per-topic catalogue and picks a file before it is allowed to open anything, instead of searching the corpus directly.
The pattern, in three layers.
Andrej Karpathy posted the shape in April 2026 and it spread quickly.[1] Instead of using a model mainly to write code, point it at a folder of markdown notes and let it maintain a persistent, interlinked wiki that gets better as new sources arrive. He reports running his own at over 100 articles and more than 400,000 words. The document is deliberately abstract about mechanics, which is worth knowing before you try to build from it: it communicates a pattern, not a spec.
The parts that matter are boringly simple, and the simplicity is load-bearing.
Everything above is the easy half. A model will happily generate four hundred thousand words of interlinked markdown, and the moment it has, you own a corpus too large to read and too small to justify a search stack. That is the position most people are actually in when they ask us whether this was worth doing.
What retrieval actually costs.
Our corpus at the time of the run held roughly 330 topic files with about 300 one-line index entries pointing at them. That is nowhere near the size that needs a vector database, which is exactly why the question is interesting: at that size, what does the agent actually spend to answer a question, and does structure beat just letting it look?
We ran two arms on six real workspace questions with file-verifiable answers. The baseline arm got free rein: grep, glob, read, whatever it wanted. The index-first arm ran one deterministic path before anything else, a 314-line script with no embeddings and no model calls anywhere in it. Strip the question to keywords, score every index line by inverse document frequency without opening a single file, open the one best file, pull the best section out of it, follow at most one pointer. Then answer from that. Marginal tokens are measured against a do-nothing control agent, so what you see below is the cost of the retrieval, not the cost of existing.
Read the shape, not the total. The 9.5% saving is the least interesting number on the chart, and on wall time the structured path was marginally slower. What the chart actually shows is two different distributions. Free search is bimodal: cheap on the four questions where the auto-loaded index head happened to name the right file, and then 33k and four tool calls on the two where it had to hunt. Index-first showed no second mode across these six. Every question, hunt-class or not, landed between 24,371 and 25,923 tokens and almost always in one tool call. On the two hunt-class questions, and two questions is all we have, the saving was over 25%.
That distinction matters more than it sounds, because the expensive failure in agent work is almost never the average case. It is the run that wanders, opens eleven files, fills its context with near-misses and then answers from a degraded window. A ceiling on retrieval cost is worth a few hundred tokens on the easy questions, and that is close to what it cost here: index-first ran a few hundred heavier on three of the six, and its median still landed slightly below free search.
The index is the asset. The script is a lens.
Here is the part we got wrong first, and it is the part worth stealing. Rounds one and two of this benchmark produced retrieval misses. Every single one traced back to three index lines, and none of them was a scoring bug. The lines were written in the vocabulary of the solution, and questions arrive in the vocabulary of the symptom. A human skims past that mismatch. A lexical scorer cannot.
The second-order finding is the one we did not expect. The baseline arm was already competitive, and the reason is that it had the rebuilt one-line catalogue auto-loaded into its context. After the rewrite, both arms stopped missing. We cannot put a number on that: the pre-rewrite index was already gone from disk, so rounds 1 and 2 against round 3 are a before-and-after across rounds, not a controlled comparison. The script is a cheap lens for the long tail past whatever your context budget auto-loads. If you only do one thing from this note, write better index lines. You can add the scoring later. We did not measure how the value splits between the index rewrite and the scoring layer, and after losing the pre-rewrite index we no longer can, so treat the ordering as our judgment rather than a result.
One limit of a lexical scorer is worth stating plainly, because it is the case our own audience hits first. IDF scoring matches words, not meanings. A question asked in French will not score an index line written in English, and a question built on a synonym of the index line's key term will not score it either. If your corpus or your team is bilingual, either write the index line in both languages or accept that a share of your questions start from a miss. That is the first real argument for embeddings, and it arrives well before corpus size becomes one.
One reading of the numbers should be resisted. We are not claiming a 9.5% improvement over a naive setup. Our own caveat at measurement time was that the true pre-rewrite baseline no longer existed on disk, because the old 99KB index had already been replaced. The measured gap understates what the rewrite was worth and says nothing about what you would see starting from an unstructured pile of notes.
Three corpora, three sets of problems.
What carries across all three is the interesting part.
The production system at the top of that ladder is Garry Tan's gbrain, and his own note on it is the most useful sentence in the repository: none of the ideas are novel, the contribution is shipping all of them together.[4] Four of its conventions cost nothing to copy and we did. A notability gate, where an entity gets a page on its second mention and not its first, because a junk page wastes attention and degrades search for everything else. A split between rewritten synthesis at the top of a page and append-only dated evidence below it, which makes staleness mechanical instead of a judgment call. Timeline merge, where an event touching three entities lands on all three pages. And gap analysis as a required part of every answer, so the system states what it does not know instead of papering over it.
From the adopter thread on the original gist, two reports are worth repeating with their tier attached. Both are self-reported and neither is verifiable.[2] The most-cited adopter, running about 4,000 pages over six months, says drift from under-updated cross-references is the headline failure mode and that the lint pass, in their words, is not optional and runs on a timer. The other is the only controlled comparison anyone in the thread claims to have run: pages stamped "derived, verify against source" made their wiki net-negative, because the agent then reads the wiki and the source and pays twice. Their conclusion, which we think is right and which cuts against how most people would hedge: either the wiki is trusted at query time or it should not exist.
There is a corollary that saves real money. If a page is a mirror of a small file the agent could already open, it is worth less than nothing. The pattern pays exactly when a page compresses facts scattered across many sources into one place. We found the same thing from the other end: agents read documents in runs and rarely follow cross-references, so a self-contained page beats a well-linked one.
How to start one without wasting a quarter.
- Write the index before you write the pages.
One line per topic, in the words a future question will use. Symptoms, error text, the thing that broke, the name someone will actually type. Not the name of the fix. This is the asset, and it is the only part a scoring function can see.
do · one line per topic, in symptom vocabulary - Start with no retrieval infrastructure at all.
At a few hundred pages the auto-loaded catalogue plus plain file reads is a serious baseline, and ours was hard to beat. Add lexical scoring when the corpus outgrows what you can afford to keep in context, and consider a vector store when it outgrows the lexical scorer. We would expect most business corpora never to reach that third step, though that is a prediction and not something we have measured across a sample.
do · earn each layer with a measurement - Decide, in writing, that the wiki is trusted at query time.
A page the agent has to double-check against a source is worse than no page. The evidence for that is one unverifiable adopter report plus our own reading of the mechanics, so weigh it accordingly. If you accept it, it means a real editorial standard on what gets written and a real lint cadence, not a hedge stamped in the frontmatter.
do · no "verify against source" escape hatches - Put the lint on a schedule, not on good intentions.
Broken links, orphan pages, contradictions between pages, claims that have gone stale. This is the maintenance humans abandon, which is the entire reason the pattern works when a model does it. Abandoning it by another route puts you back where you started.
do · scheduled lint, results reviewed by a person - Do not file every good answer back into the corpus.
The original pattern suggests it and the practitioners who have run one longest push back. Derived pages that add no new information silt up the corpus and make every future search harder. Apply a notability gate to synthesis the same way you apply one to entities.
do · a page must compress something, or it does not exist - Measure your own worst case before you trust it.
Six questions with verifiable answers, two arms, count the tokens and the tool calls. That is an afternoon. You are looking for one thing: whether the expensive tail exists at all.
do · benchmark the tail, not the median
What this does not license.
Six questions is six questions. One corpus, one harness, one run per arm, and the people who built the thing being tested are the people who measured it. Our token cells are solid and our timing cells are indicative at best, with several seconds of noise per run, which is why we are not claiming the speed result in either direction. Anyone reporting this as a general finding about retrieval architectures is over-reading it by a wide margin. What it supports is narrower and still useful: on this class of corpus, a deterministic index-first path removed the expensive tail without costing correctness.
Two things make us take the shape more seriously than a single internal run deserves on its own. A group at Tencent and several universities reported a result in July with the same geometry on a completely different substrate, which we have not reproduced: a maintained, behaviour-centric index navigated before free repository exploration improved edit-site localization while cutting planner tokens, with the biggest gains on the search-hostile cases.[3] That is the codebase analogue of our hunt-class questions. And there is a blunter corroboration from a shipped tool: repomix exists to pack an entire repository into one file for an LLM, and its own documentation tells agents not to load that file, to retrieve from it incrementally instead.[6] That is a documentation recommendation rather than a measurement, and we have not benchmarked it. We read it as a tool author arriving at the same conclusion from the opposite end.
The corollary cuts at us too, and it is worth saying out loud. If your agent already has read, grep and glob over a local folder, packing that folder into a context-filling artifact first is the arm this benchmark says loses. Cheap retrieval is not a feature you add. It is a thing you stop preventing.
If you are sitting on a pile of internal documents and wondering whether an AI-maintained knowledge base would pay for itself, the honest first step is a measurement. Build after that, if the measurement says to. The contact form reaches us, and we will send back a written read on what your corpus would actually cost to search, free. We do not need your documents to do it: a file count, a size, and a sample of the questions people actually ask is enough. If an assessment does eventually need to touch your content, that is a Law 25 conversation before it is a technical one.