← Journal

What Does Your System Know When It Starts Thinking?

Everybody is selling AI. The product may actually be context. Some of it gets retrieved, some gets assembled at the moment of the call, and some of it gets compiled into small models, which is how I ended up with a machine that had my accent and none of my facts. On context compilers, working sets, the rise of the micro model, whose weights they are anyway, and what to do about it on Monday.

A vast dark archive at night, shelves of files, ledgers and server racks receding into blackness on both sides. A lone figure in a dark coat stands with his back to us at a small desk, one hand on a single open envelope. Of the thousands of documents on the shelves, only a few dozen glow electric blue, and thin luminous threads run from each of them across the dark and converge on the envelope, where they meet in a single hot coral spark. Everything else stays unlit.

Context is becoming a first-class software asset. The model is becoming the thing that executes it.

This is a long one, in three parts. The first is the argument. The second is what it means for anyone choosing technology. The third is how to actually do it. Read the first and come back for the rest when you have a decision to make.

Who this is for · jump to your part
Everyone

The argument, in one sitting. Start at three models, one prompt and read to what it actually sells.Part one · about fifteen minutes

Boards and executives

Why the room is out of context, and the five things it needs before it decides. The board is out of context too, then changing the room's mind.Part two · about ten minutes · two diagrams

CEOs, CTOs and strategy

The eight technology decisions and the steering-committee table of old questions against new. The decisions.Part two · about seven minutes

Engineering leadership

The method, in the order to do it now: one decision, four bins, the evaluation set, the envelope assembler, retrieval, the tiered router, the compile step, and the five numbers. How to do it.Part three · about twelve minutes · one architecture diagram

Legal and commercial

Once a customer's knowledge has been compiled into a model, whose asset are the weights? Whose weights are they?Part one · three minutes · the clause nobody has written yet

In a hurry

The answer to the title is at the end. So, what does it know? It will make more sense if you read the rest, but it stands on its own.Two minutes

Part one — The argument

“Approve it.”

On its own, that sentence means almost nothing. Approve what? Who is asking? Under whose authority? For which customer, against which contract, after what happened last time, and what exactly are we trying to achieve by approving it?

Humans resolve all of that without noticing. We walk into the room already carrying the meeting before, the client’s temper, the number the board will not go above and the name of the person who will be blamed if it goes wrong. We call that context and we never think about where it came from.

A language model walks into the room with none of it. It has read most of the internet and it knows nothing about your situation. Every call starts from a blank desk.

For the last couple of years the industry has been arguing about which model is cleverest. I think that argument is about to become the least interesting one in the room. The interesting one is what the system knows at the moment the model starts thinking, and who built the machinery that put it there.

I have been circling this for a while

Before writing this I did something slightly vain and searched my own journal for the word. Fifty-five essays as of this week. The word “context” appears in twenty of them, forty-six times in all, and the shape of where it turns up is more interesting than the count.

PublishedEssayMentions
1 Oct 2025Why Imperfection Is the Future of Intelligence2
4 Mar 2026What Seven Years Inside a Construction Giant Taught Me About Risk, Scale and Survival2
12 May 2026The Future Has No Target State2
20 May 2026The CTO Is the First Executive Role to Break1
2 Jul 2026The Senses Were Always the First Interface2
4 Jul 2026Trust Is the New Infrastructure1
7 Jul 2026Why Philosophy Is the Backbone of Critical Thinking in the Age of AI2
8 Jul 2026The J-Space: On Catching a Thought Before It Speaks2
8 Jul 2026The Paradox of the Empty Pocket2
13 Jul 2026Move Fast and Break Compliance2
13 Jul 2026Off Topic and Personal: The Annoyance and Delight of Being Human1
13 Jul 2026The Prediction Decade: Why Seeing the Future Just Became the Most Valuable Thing You Can Do2
13 Jul 2026The Things I Learnt Building Agencio Predict’s Signal Fabric — and Where They Travel1
16 Jul 2026The Governor — Why the Agentic Age Needs an Operating System1
17 Jul 2026The Economics of Thought12
22 Jul 2026Understanding Marketing — Behavioural Science, Data, and the Machine I Built to Serve It1
6 Aug 2026Twenty Years On: Why Social Media Needs a Reboot1
12 Aug 2026The Invisible Ink Problem4
9 Sep 2026A Brief History of $317 Billion1
17 Sep 2026When AI Has Something to Lose4

It also appears ten times across the product pages, which is where it was doing its real work all along.

The first mention is from October 2025, in an essay about why imperfect intelligence is the useful kind, where it was a supporting word and not the subject. It stays in the margins for the better part of a year, one or two appearances per piece, always in service of something else: risk, trust, compliance, the senses, the CTO breaking. Then in July 2026 it moves to the centre. The Economics of Thought uses it a dozen times, because that was the essay where I noticed that better than ninety-nine per cent of a day’s machine effort was not thinking but remembering, and remembering is what context costs when you have not decided where to keep it.

So I have been talking around this for a year without once giving it a title. That is usually a sign that the thing is real and the vocabulary was late. This is the essay where it gets the title.

Three models, one prompt

Try an experiment. Take three capable frontier models and give them the same generic prompt. Draft a response to this tender. Summarise this contract. Tell me whether this risk is acceptable.

The answers will be competent, fluent, and much the same. Anyone who has run this comparison in anger knows the slightly deflating feeling of realising the models have converged and the difference is now mostly price and latency.

Now give one of the three something the others do not have. The buyer’s history. The last six bids and why each one was lost. The engineering evidence the firm actually holds. The margin floor. The two people internally who understand the missing discipline. The board’s definition of unacceptable risk, in the board’s own words, from the minutes.

That one is a different product. Same model, better context system. A generic model knows what a tender is. A contextualised system knows this tender, this buyer, and what happened the last six times. The gap between those two is where the value is.

What we actually ship now

For thirty years software companies shipped functionality. Code that did things. Then AI arrived and, for a brief and slightly giddy period, we shipped intelligence, which mostly meant a text box connected to somebody else’s API.

What ships now, if you look at any AI product that works, is a bundle. Model, tools, memory, knowledge, identity, state, rules, workflow. The model is the least proprietary part of that list.

Every time the system performs a task, the software assembles an envelope around the model. Here is who the user is. Here is what they are trying to do. Here is what happened before. Here is the evidence. Here are the rules and the tools and the things you must not do. Here is the current state of the world. Then the model reasons inside it.

The envelope is the product. The model is the thing that executes it.

An infographic titled Context. The most important word in AI, and the most misunderstood. Four panels. The confusion around context, a cloud of overlapping bubbles: a bigger prompt, my documents, the internet, a longer memory, real-time data, the model's knowledge, RAG, tokens, fine-tuning, previous chats, user permissions, my company data, system state, all of the above. What context really is, a stack of six layers: user and identity, goal and task, knowledge, current state, tools and actions, rules and guardrails, resolving to a coherent minimal context envelope for the task, not everything, just what is needed. AI plus human judgement, a Venn diagram: AI processes, reasons at speed, stays consistent, never tires; the human understands nuance, brings experience, sees the bigger picture, handles ambiguity, applies values, takes responsibility; the overlap is better questions, better answers, better decisions. Why it is valuable: higher quality decisions, lower cost, lower risk, greater productivity, compounding value. The footer reads Context is not a buzzword, it is a competitive advantage.

Everyone uses the word. It means a different thing to each of them, which is the first problem the compiler has to solve.

Where context lives

If context is an asset, the first design question is where to keep it. There are four answers, and the difference between a system that works and one that is merely expensive is mostly knowing which answer applies.

In the weights. A small specialist model can internalise stable patterns. Terminology. Classification schemes. House style. Domain conventions. The way this firm scores a risk. You pay for this once, at training, and it is close to free at every call afterwards.

In knowledge. Documents, databases, previous cases, contracts, bid libraries, specifications, research. Too large or too changeable to bake into a model. This is retrieved when needed.

In state. What is happening right now. This user, this customer, this transaction, this tender, this agent run, what the other agents did an hour ago. This is assembled fresh for every call and thrown away afterwards.

In orchestration. Which model should see what. Which tools it may use. What it should remember and what it should discard. Where the result goes next. This is the layer that is bigger than retrieval or prompt engineering, and the one most products have not built yet. It is the layer I sketched as an operating system for agents in July, before I had a name for what it was compiling.

The rule that sorts things between them is stability. Volatile knowledge in the weights goes stale and starts confidently telling you about last year’s org chart. Stable knowledge in the prompt gets paid for on every single call, forever, which is the software equivalent of re-explaining the company to a new contractor every morning.

If the essay has a governing rule, it is this one, and everything after it is an argument for taking it seriously.

Compile what is stable. Retrieve what is large. Inject what is current.

Orchestration is the part that decides which is which.

Where context lives · compile what is stable, retrieve what is large, inject what is current
BinWhat belongs thereHow it is paid for, and how it fails
Weightscompiled

Stable and hot. Terminology, taxonomies, house style, domain conventions, the firm's way of scoring a risk. Distilled routines.

Once, at training. Near free per call. Fails by going stale, and by rounding away if you compile it too hard.

Knowledgeretrieved

Large or slowly changing. Documents, cases, contracts, bid libraries, specifications, research, with provenance and a date.

Per call, for what is fetched. Fails by fetching too much, and by trusting a chunk nobody has touched since the reorganisation.

Stateassembled

Volatile. This user, this customer, this tender, this agent run, what the other agents did an hour ago.

Per call, every call, then thrown away. Fails by arriving late or filtered, and by being paid for as if it were stable.

Orchestrationthe linker

The rules of the envelope. Who sees what, which model, which tools, what is remembered, what is discarded, where the result goes.

Once, in code you own. Fails when it is outsourced, or when it is an agenda set by whoever has the proposal.

The lifecycle. An expert's correction starts in state → kept, it becomes knowledge → stable and hot for long enough, and proven by the tests, it is compiled into the weights. That promotion pipeline, and the evaluation set behind it, is the moat.

I have a number for that, from an earlier essay. In one long agentic session I audited, the system read 1.3 billion tokens from cache against 2.4 million read fresh. For every token it learnt, it re-read something it already knew roughly five hundred times. The caching is what made the day affordable. It is also a measure of how much of the work was simply describing the situation again.

The context compiler

Here is the idea I think is at the centre of all this.

Today’s applications take data and execute predetermined logic against it. The next generation takes an enormous quantity of messy information and compiles the relevant parts into context for an intelligent system. The interface says “help me respond to this tender”. Behind it, the application constructs the world the model needs to understand before it answers, down to who internally knows the thing nobody wrote down.

That application is not calling a language model. It is assembling reality for one. I have started calling it a context compiler, because the analogy holds further than I expected.

A compiler decides what to resolve at compile time and what to link at run time. Static linking bakes the library into the binary. Dynamic linking looks it up when the program runs. Neither is right; the choice depends on how often the thing changes and how often it is used. That is exactly the decision the four locations above force on you. The weights are static linking. Retrieval and state are dynamic. Orchestration is the linker.

It goes further, because context has a lifecycle. Something starts in state: an expert corrects the machine once, in a session. If the correction is kept, it becomes knowledge and gets retrieved next time. Once it has proven stable, and is being retrieved often enough, it is worth compiling into a specialist model, where it costs nothing at all per call. That is a just-in-time compiler. Interpret everything, watch for the hot paths, compile the ones that run constantly. The context compiler is a JIT for organisational knowledge.

Which is the whole thesis in one sentence, if you want it that way. The system’s job is to decide what each form of intelligence should know, where that knowledge should live, and when it should be promoted from state to knowledge to weights.

An infographic titled Just-in-Time LLM Training Pipeline, the right intelligence, compiled when needed, ready to use. Six numbered stages across the top. One, detect need: the system realises it repeatedly needs a new capability, from user requests, recurring tasks, gaps in answers, low confidence, or high cost and latency, with a trigger threshold. Two, assemble data: internal documents, previous conversations, domain data, tool outputs and logs, human examples; clean, chunk, label, deduplicate, apply permissions and filters. Three, train, automated: prepare the training set, fine-tune or distil, evaluate, optimise by quantising or pruning, as CI/CD for models, reproducible, versioned, logged. Four, validate: quality benchmarks, domain test set, guardrails and safety, bias and risk checks, cost and latency targets, comparison with existing, human spot check, then approved for deployment with a model card. Five, package and deploy: containerise, register in the model catalogue, deploy to edge, cloud or local, expose via an internal API. Six, use and improve: route matching requests to the new model, monitor, collect feedback, retrain or update, retire when no longer useful. Below, the runtime decision flow: a user request, analyse this tender against our criteria, goes to the context compiler, which retrieves relevant context, selects the best models, applies rules and constraints and builds the model input, the envelope; a decision diamond asks whether a specialist model can do this; yes runs the specialist small model, fast and low cost; no uses the frontier model for complex or novel tasks; either way the answer comes back with sources and a confidence figure. A dotted line marked new capability required loops back to stage one. A feedback loop on the right logs interactions, identifies new training opportunities, updates datasets and retrains, evolves the model portfolio and deprecates unused models. Key benefits along the bottom: on-demand capability, lower cost, faster and more relevant answers, domain expertise compiled rather than prompted, controlled and auditable, continuously improving.

The promotion pipeline, drawn as a pipeline. The dotted line at the bottom, new capability required, is the hot path being noticed. Everything above it is the compile.

Which tells you where the moat is. Not in the documents, which a competitor can buy or scrape. In the promotion pipeline, and in the evaluation set that proves each promotion held. Two firms can use the same model and the same document estate and get radically different results, because one of them has spent five years learning what is stable enough to compile and has the test results to prove it.

I should admit a bias here. I spent a stretch of the late nineties and early two-thousands on a processor whose entire premise was that the compiler would be clever enough to find the parallelism. It was not always. So I say “the compiler will sort it out” with the appropriate humility. But the shape of the argument is right, and I have the scar tissue to know which bits are hard.

The dragon, and the flag I set too high

If you want proof that context can be compiled into a model, I did it on the sofa last week.

A seven-billion-parameter Qwen, a rank-eight adapter trained on about three hundred and thirty thousand words of my own writing, one laptop, six hours. About 0.15 per cent of the animal was trainable. At the end of it I had a model that wrote in my register and, after a second run with better food, could tell you what my products were and who they were for, from the name alone. Stable knowledge. Compiled. Nearly free to run.

Then I tried to ship it small.

The obvious export is to fuse the adapter into the four-bit base and keep the file at four gigabytes. I did that, and asked it about Silo, and it told me with complete confidence that Silo was a content management system with Markdown templating. Silo is a security architecture for watching AI agents. The voice was still perfectly mine. The facts had gone.

The arithmetic is unforgiving. A light adapter moves each weight by less than one four-bit quantisation step. Re-quantise after the fuse and most of the change rounds back to where it started. The voice is spread across millions of weights and shrugs that off. The memory of one product page lives in a handful of weights and vanishes. At eight bits the step is sixteen times smaller and everything survives.

So I compiled my own organisation into a machine, set the optimisation flag one notch too high, and got back something with my accent and none of my facts. Anyone who has shipped a release build with the wrong flags will recognise the feeling.

It is also the useful lesson. Compiled context is real, it is cheap, and it is precision-sensitive in ways nobody warns you about. Compile too hard and the facts get optimised away. Every compiler has flags. So does this one.

The working set

There is a tempting counter-argument, and it comes with a million-token context window attached.

Why bother deciding what belongs where, when you can give the model everything? Windows are enormous now. Put the whole document estate in the prompt and let the model sort it out.

Computer science answered this in 1968. Peter Denning’s working set model says a running program needs only a small, shifting subset of its memory resident at any one time. Give it that subset and it runs. Page in far more than it needs and the machine spends its effort shuffling memory instead of computing, a condition Denning named thrashing, and which anyone who used a laptop with too little RAM has heard through the fan.

Long context has its own version of thrashing. The research literature calls it lost in the middle: models attend well to the start and end of a huge prompt and poorly to the bulk of it. Cost scales with what you put in. Attention does not. Being able to give a model everything does not mean you should. Expertise is partly knowing what not to consider. The sophistication of a context system is not measured by how much it can stuff into a window. It is measured by how reliably it finds the minimum sufficient context for the decision in front of it. The working set, in Denning’s phrase, fifty-eight years on.

The future of AI may not belong to the model that knows the most. It may belong to the system that knows what the model needs to know next.

The stack

None of this makes the frontier model go away. It moves it up.

The assumption that a bigger model makes a better product does not survive contact with the bill. There are tasks where frontier intelligence genuinely earns its cost: novel reasoning, ambiguity, planning, synthesis, difficult judgement, deciding what to do when the plan breaks. And there are a great many tasks where it is an absurd quantity of machinery for the problem. Classify this. Extract that. Route this to the right queue. Score this against the rubric. Say it in the house style.

A small model adapted for one of those jobs is faster, cheaper, more predictable and, the part that gets underrated, runs inside the building. For a bank in Singapore or anyone with a security clearance, the data never leaving is the sale, not the saving.

One refinement to the usual pitch, because “a cheaper AI” undersells what a specialist model is. It is stable knowledge and stable behaviour compiled into a reusable cognitive component: the first line of the rule above, made into a thing you can version, test and ship. Small models are vastly cheaper to train from scratch than frontier ones, but that is not the commercially important point, because almost nobody should be training a foundation model. The important point is that they are cheaper to adapt and cheaper to run. You take someone else’s dragon and put your saddle on it. And be careful with the claim that a small model can learn your reasoning. It can learn to reproduce a routine on familiar inputs, which is distillation, and it is very useful. It is not the same as thinking, and the sceptic in the room will go straight for that sentence if you blur it.

So the application starts to look like a computer rather than a chatbot. The frontier model is the processor you reach for on the hard branches. The micro models are the fixed-function units, the bits of silicon that do one transform very fast. Retrieval is main memory. The archive is disk. The context window is the cache, and the whole art is cache management. Deterministic code does what deterministic code has always been good at, which is being right the same way every time.

Code where certainty matters. Micro models where specialisation matters. Frontier models where intelligence matters. Context everywhere. And above all of it a person who answers for the consequence, which is where the scarcity actually sits, as I argued in The Supervisory Economy.

An infographic titled Many minds. One outcome. Smaller models bring specific context, larger models bring broader reasoning. Four columns left to right. Your world: documents, data, people, tools, experience, external. Specialist small models: a document model, a domain model, an analysis model, a compliance model, an agent model, each fine-tuned or adapted on your own data. Context orchestration: selects what is relevant, combines insights, manages memory and state, applies permissions, controls cost and latency, decides when to use a larger model. Frontier model, when needed: deep reasoning, creative synthesis, ambiguity, connections, used for complex reasoning, novel situations, cross-domain thinking, strategic judgement and open-ended answers. A side panel lists why small models are valuable, how to create them in four steps, and when a larger model may not be needed. The footer reads The right model, the right context, a smarter outcome.

The stack, drawn out. The frontier model is on the right, where it belongs, and it is the smallest box.

I have three products that sit at different points of this, and I will name them once and move on. CodeEasy is a context compiler for a codebase, and has a tool in its interface called, without much imagination, get_context. The jjvoice model is the micro model, compiled from this journal. The Governor is the orchestration layer, the thing that decides which model gets which envelope and what it is allowed to do with it. I did not set out to build the stack in this essay. It turned out that way, which is either validation or a warning about the author.

Whose weights are they?

There is a commercial question underneath all this that I have not seen anybody write down yet, so I will.

Suppose a vertical software company does the sensible thing. It takes a customer’s twenty years of engineering judgement, compiles the stable parts into a specialist model, and ships that model as part of the product. The customer’s decisions get better. Everyone is pleased.

Whose asset is the model?

The training data was theirs. The compiler, the pipeline, the evaluation set and the eight-bit lesson were ours. The weights are a blend of the two that nobody can unpick. When the contract ends, or the vendor is acquired, or a competitor becomes a customer of the same vendor, that question stops being philosophical.

Source code had this fight decades ago and settled it with licences. Data had it more recently and settled it, roughly, with residency clauses. Compiled context is going to have it next, and the firms that put a sensible answer in their terms now will look considerably more grown-up than the ones that discover the problem in a dispute.

What it actually sells

Customers do not buy context. They would look at you oddly if you offered it.

They buy an answer to “should we bid for this?” from a machine that understands what this means, and that knows which of everything it could know is actually relevant this morning, and leaves the rest on disk where it belongs.

That is not a model. It is a decision environment. I suspect it is what the next generation of vertical software companies actually sells, whatever they put on the website.

Which means the buyer’s question changes, from “what model do you use?” to the one at the top of this page.

Part two — What this means on Monday

The board is out of context too

Before the decisions, a word about who makes them.

Everything in part one was about what a machine knows at the moment it acts. The uncomfortable observation is that the same test applies to the room where the technology budget gets signed, a room whose technology seat has already fractured under the same pressure. Ask what the board knows at the moment it decides about AI, and the honest answer is usually: a vocabulary picked up from the press, a pack read on the plane, a management summary that is three weeks old, and a great deal of confidence.

That is not a criticism of boards. It is the same architecture problem wearing a suit. Twenty years of pattern-matching is knowledge compiled in an earlier release of the world. A two-hundred-page pack is thrashing in a folder. State arrives filtered through management. And the orchestration is the agenda, set by the people whose proposal is on it.

The perception of understanding is the dangerous part. A board that knows it does not understand asks good questions. A board that has heard the words often enough to feel fluent approves the wrong thing with the same confidence a four-bit model brings to describing a product it has forgotten. I have sat in both rooms. The second one is quieter and more expensive, and it fails the way every room that cannot be told it is wrong fails: by removing the people who would have said so.

So the board needs an envelope too. Not a course, and not a demo, both of which produce fluency without context. The smallest sufficient representation of reality, assembled for the decision in front of them.

An infographic titled The board is a decision-maker with a context problem. Same words, different outcome, it depends on what the room knows. On the left, the world, everything that exists: documents, data, people, systems, history, external. An arrow marked retrieval, find what matters, leads to the board room, a smaller more useful view of reality, drawn as six figures around a table. An arrow marked orchestration, decides what reaches the room, leads to the outcome today: a speech bubble saying Looks good, let's approve it, with five crosses beneath: on the wrong data, missing key context, important history ignored, right questions not asked, high confidence, wrong outcome. Below the room sits the board envelope, the smallest sufficient representation of reality, with five items: weights, what it already knows; retrieval, what we bring in; state, who, what, now; orchestration, what it sees and can do; the agenda, what reaches the room. Along the bottom, what the room needs when it decides, five cards: the vocabulary, the mechanism, the consequence, the decision, the questions. The footer reads Confidence is not context.

The same four places the machine fails, drawn for the room that signs the budget.

The right-hand column is the elevation. It goes in that order because each step earns the next. Vocabulary without mechanism is jargon. Mechanism without consequence is a lecture. Consequence without a decision is a strategy offsite. And the questions at the end are the part that lasts, because a board that knows what to ask no longer depends on being told.

The point of doing this is not to turn directors into engineers. It is to put the decision in front of people who are in context when they make it, which is the only condition under which a business can actually pivot rather than merely announce that it has. Most of the failed AI programmes I have seen were not failures of technology. They were approved by a room that was out of context and did not know it, which is the one failure mode the machine and the board have in common.

Changing the room’s mind

Knowing the board is out of context is the diagnosis. The harder question is what to do about it, and the honest answer starts with what does not work.

Arguing does not work. Nobody was ever argued out of a belief they arrived at by feeling fluent. Courses produce more vocabulary, which is the thing that caused the problem. Demos are worse, since a good demo manufactures exactly the sensation of understanding the room already has too much of. What changes a mind is a decision, with a deadline, made with better context than last time, and a result the person can see.

So the plan has the same shape as the compiler. Find out what the room knows. Remove the false fluency. Put the smallest sufficient context in front of them. Give them a real decision. Measure it. Then, and this is the part people find hardest, accept the result either way.

An infographic titled Changing the room's mind. One board cycle, then decide. Five numbered steps arranged in a circle around a boardroom table. One, diagnose which room you are in, listen to the questions before you say anything; the signal is that the questions change before the answers do. Two, combat the false fluency, not by arguing but by one small, safe, visible failure in their own domain; the signal is somebody saying I thought it knew that. Three, educate with the envelope, not a course, ninety minutes once a cycle, vocabulary, mechanism, consequence, decision, questions; the signal is a director asking a vendor the title of this essay unprompted. Four, give them a real decision, small, reversible, with a number attached and a quarter to prove it; boards move for a peer, not a supplier; the signal is a decision with a name on it and a date. Five, measure it in their numbers, cost per decision, the working set, the pass rate, in front of the same room one cycle later; the signal is the room asking for the next one. Below, two outcomes. The mind changed: they become the proof for the next room. It did not, after two cycles: shake hands and move on, not every business needs to pivot this year, other people need the help. The footer reads Then the next room. The proof travels with you.

One cycle. Five steps, each with the signal that it worked, and a fork at the end that is a kindness, not a judgement.

Two things about the fork, because it is where most advisers go wrong.

The first is that the time-box is a kindness, not a judgement. A board that will not change is not stupid. It has usually decided, without saying so, that the cost of being wrong later is lower than the cost of moving now, and sometimes it is right. Two cycles is enough to know. Staying longer is not persistence, it is a consultant’s sunk cost, and it deprives a room down the road that was ready and waiting.

The second is that the ones who changed are the whole engine. I have never persuaded a board with a slide. I have watched boards persuade each other in a corridor with a number, in under a minute. The plan above exists to manufacture that number and the person willing to say it out loud. After that, the mindset changes itself, one room at a time, and your job is mostly to make sure the proof travels.

The decisions

All of which is pleasant to think about and useless unless it changes a decision. So here is where I think it lands for anyone choosing technology for the next few years, whether from the boardroom or the engineering floor.

Stop procuring models. Procure the ability to swap them. The model is a commodity input, like compute. Every vendor conversation should establish how you change model, how long it takes, and what breaks. If the answer is “you can’t” or “everything”, you are buying a model with a product attached, and the model is the part that will be obsolete in eighteen months. The line that matters is the cost to change, not the cost to buy.

Rent the dragon. Own the saddle. The orchestration layer, the thing that decides what each model sees and may do, is the one part of the stack you should not outsource. It is where your judgement about your own business gets encoded. A firm that rents its models and owns its context compiler is in a strong position. A firm that does the reverse has built its competitive advantage on somebody else’s roadmap.

Audit where your context lives now. Most organisations, asked this honestly, discover that the stable knowledge is in six people’s heads, the large knowledge is in a SharePoint nobody trusts, and the state is in email. Inventory it. Sort it by the stability test. What is stable enough to compile, what is large and slow enough to retrieve, what is volatile and must be assembled fresh. That inventory is the first real architecture document of the AI era, and hardly anyone has one. Unlike a target state, it is allowed to change every quarter, and it should.

Build the evaluation set before you build anything clever. You cannot promote knowledge into weights, or even into retrieval, unless you can prove the promotion held. My four-bit fuse looked fine until somebody asked it a question with a known answer. The test set is the asset. Fund it first, keep it current, and treat a promotion without a passing test the way you would treat a release without one.

Budget per decision, not per token. Cost per token rewards stuffing the window. Cost per decision rewards finding the working set. Measure how much of each call is re-describing the organisation to the model, and treat that number the way you would treat any other recurring cost that could be capitalised instead.

Decide sovereignty before cost. A small model that runs inside the building is a different category of decision from a cheaper one that does not. For regulated data, defence, health, or anything a client would be alarmed to find in a third-party log, the micro model is not the economy option. Trust is the infrastructure here, and it is cheaper to build in than to bolt on. It is the only option, and the frontier model gets the parts of the problem that can leave.

Put the weights clause in the contract now. On whichever side of the table you sit. Who owns a model compiled from your data with their pipeline, what happens to it at exit, and whether it can be used to serve your competitor. Ten minutes with a lawyer this year saves a year with several of them later.

Change the vendor question. Ask what their system knows about you when it starts thinking, how that knowledge got there, and what happens to it when you leave. Vendors who can answer that crisply have built a context system. Vendors who reach for a slide about their model partnership have built a text box.

If you prefer it as a table for the next steering committee, it comes down to this.

The question you used to askThe question to ask now
Which model do you use?What does your system know about us when it starts thinking?
How big is the context window?How do you find the working set for a decision?
Can it read all our documents?Which of our knowledge is stable enough to compile, and how do you prove it held?
How much per token?How much per decision, and how much of that is re-describing us?
Is it on-premise?Which parts of the problem never leave the building, and which model handles them?
What is the licence?Who owns the weights trained on our data, and what happens at exit?

None of this requires a new model. It requires deciding what belongs where, which is a strategy problem wearing an engineering costume, and has been all along. Strategy is becoming software; this is what the software is for.

Part three — How to do it

Strategy is cheap to write and expensive to be wrong about, so this last part is the engineering. It is the order I would do it in now, which is not the order I did it in, and the difference is most of what I learnt.

Start with one decision, not with the data

Every context project I have seen fail started with the corpus. Ingest everything, embed everything, then look for a use. That is building a library and hoping somebody will need a book.

Start instead with one decision the business makes repeatedly and expensively. Should we bid. Is this claim valid. Which supplier. Does this design pass review. Then sit with the best person who makes that decision and write down what they know at the moment they make it. Not what they could look up. What is actually in their head and on their desk. That list is the context specification, and it is usually a page long, which tells you how small a working set really is.

Sort it into the four bins

Take every item on that page and ask three questions. How often does it change. How often is it used. What happens if it is wrong.

Then apply the rule. Compile what is stable. Retrieve what is large. Inject what is current. Stable and hot goes to the weights, eventually. Large and slow goes to retrieval. Volatile goes to state and is assembled fresh. The remainder, the rules about who may see what and which tool does what, is orchestration. Write the bin against each item. Argue about the borderline cases, because the argument is where the architecture gets decided, and it is cheaper to have it on a whiteboard than in production.

You will discover that the stable, hot knowledge lives in six people’s heads, three of whom are on leave. That is not a blocker. It is the project.

Build the evaluation set before anything else

This is the step everybody skips and the one I would now fund first.

Write a few hundred questions with known answers, drawn from the decision. Some should be answerable only from the stable knowledge, some only from retrieval, some only from state. Ask about things by name alone, with no description in the question, because that is how real people ask and it is exactly what my first dragon could not do. Hold a portion back that nothing is ever trained or tuned on, so that at the end there is an honest number rather than a flattering one.

This set is the asset. It is how the system stays correctable, which is the only advantage that lasts. It is how you know a retrieval change helped, how you know a compiled model kept its facts, and how you know which optimisation flag was one notch too high. Keep it in version control next to the code. Add to it every time the system gets something wrong in the field, because each of those is a free test case that has already cost you once.

Write the envelope assembler

Now the first real piece of software, and it is deliberately boring. A function that, given a decision request and the current state, assembles the envelope. Who is asking. What they are trying to do. What happened before. The evidence. The rules. The tools. The prohibitions. The state of the world. In that order, with a token budget for each section, and a log of exactly what was produced.

Two rules. It is deterministic code, not a model, because you need to be able to reproduce any envelope from its inputs. And it is inspectable: at any point you can print precisely what the model saw. When a decision is questioned six months from now, that printout is the audit record, and it is worth more than the model version.

This is the context compiler at version zero. Everything else in this part is an optimisation of it.

An infographic titled The context compiler at version zero, what an engineer actually builds. From a request to a reasoned decision: deterministic code, specialised models and frontier intelligence, all working with the right context. Five numbered bands. One, in: a decision request, should we bid for this; identity and permissions, who is asking and what can they access; session state, previous turns, current task, constraints; what other agents did. Two, assembler: the envelope assembler, version zero, printable, reproducible, with eight sections: who, objective, history, evidence cited with provenance, rules, tools, prohibitions, state of the world; rules and prohibitions outlined in coral. Three, feeds: retrieval, chunks with provenance and a date; memory, with a lifetime and a deletion path; compiled model, stable knowledge fused at eight bits; live systems, state, APIs, read fresh. Four, router: try code first, then a micro model, then escalate to the frontier model for the hard branch only, using the simplest intelligence that can do the job. Five, out: the decision with citations, the envelope logged, which tier answered and why, what was excluded and why, a decision you can trust, reproduce and improve next time. A footer: besides it, always running, the evaluation set with a held-out portion, the five weekly numbers, and the promotion pipeline that watches the logs for knowledge stable and hot enough to compile.

Version zero. Everything in this part is an optimisation of this diagram, and most of it is plumbing.

Make retrieval boring and honest

Retrieval is where projects burn their first quarter, and it is mostly because they treat it as a research problem. It is a plumbing problem.

Chunk documents along their own structure, not by character count. Attach provenance to every chunk: source, date, owner, and the version it came from. Provenance nobody can inspect is not provenance. Return citations with the text, so the envelope can say where a fact came from and the model can be asked to show its working. Give each envelope section a budget and measure how full it runs. If a section is always full, the working set is too big and you are paying to thrash. If it is always empty, the bin was wrong.

Freshness is a field, not a feeling. A chunk that has not been touched since the reorganisation should say so, and the assembler should prefer the newer one or say why it did not.

Route by tier, and log which tier answered

Orchestration starts as a router, and the router starts as an if statement.

Deterministic code first: lookups, rules, validation, anything where the right answer is the same every time. Then a small classifier for routing, extraction and scoring, cheap enough to run on everything. Then, and only then, the frontier model for the branch that needs judgement. Log which tier answered every request, because the distribution is the most useful chart you will have. The governance side of this, bounding what the agents may spend and proving they did the work, is running on my desk already. If the frontier tier is handling most of the volume, the lower tiers are not doing their job and the bill will tell you so before the chart does.

Every call that goes up a tier should carry a reason, so you can find the routines the small model should learn next.

Compile, carefully

After a few months the logs will show you knowledge that is retrieved in most envelopes, has not changed in weeks, and passes its tests every time. That is a compile candidate.

Take a small open model. Train a light adapter on the candidate knowledge, phrased as the questions people actually ask. Fuse at eight bits, not four, or the facts round away and you get my accent problem. Run the held-out evaluation, and have a model from a different vendor mark it than the one that trained it, a rule I set out in Three inputs and an adversary. Compare it against the retrieval baseline on the same questions. Promote only if it holds, and keep the adapter separate from the base so that swapping the base later is a retrain of megabytes rather than a rebuild of gigabytes.

Tie the retraining cadence to the change rate of what was compiled. Terminology, quarterly. House style, annually. The org chart, never, because it belongs in retrieval and it is only in the model because somebody was in a hurry.

None of this needs a data centre. My last run was six hours on a laptop and a pot of coffee, and the only special equipment was a command to stop the machine falling asleep.

Decide what it forgets

Memory is the part people design last and regret first. Sidekick exists because mine was failing quietly all the time, and the lesson from building it was that memory you cannot verify or delete is a liability with a search box.

Write down what persists across sessions and what does not. Give session memory a lifetime. Give customer memory an owner and an export. Give everything a deletion path that actually deletes, including from the retrieval index and from any compiled model that learnt it, which is a harder promise than it sounds and should be made honestly or not at all.

Then make forgetting visible. The envelope should be able to say “the following was excluded and why”, because a system that can explain what it chose not to consider is a system a regulator can talk to.

Measure per decision

Five numbers, on one page, refreshed weekly.

Cost per decision, all tiers included. The fraction of tokens in each envelope that are re-describing the organisation rather than carrying anything new. The working set size, as a distribution rather than an average. The evaluation pass rate, overall and per bin. And the tier distribution, so you can see the routine work moving down the stack over time, which is what success looks like on a chart.

If cost per decision is falling while pass rate holds, the compiler is working. If cost is falling and pass rate is falling, you have set a flag too high somewhere, and the eval set will tell you which.

Roughly how long

For one decision, with a small team that already knows the business, this is a quarter’s work to the first honest measurement and a second quarter to the first compile. The temptation to skip to the compile because it is the interesting bit should be resisted, for the same reason you do not skip the tests because the code looks fine.

WeeksWhat gets builtWhat proves it
1 to 2One decision chosen; the one-page context specThe best decision-maker signs it
3 to 4The four-bin inventoryEvery item has a bin and an owner
5 to 8The evaluation set; the envelope assembler at version zeroA printable envelope and a baseline pass rate
9 to 12Retrieval with provenance; the tiered routerPass rate up, cost per decision known
13 to 20First compile candidate trained, fused at eight bits, testedHeld-out pass rate holds against the retrieval baseline
OngoingForgetting rules, the five numbers, the vendor question in every contractThe chart shows routine work moving down the stack

That is the whole method. It is not glamorous, and most of it is software engineering we already knew how to do before the models arrived. That is rather the point.

So, what does it know?

01
It knows what it was compiled with.

The stable things, the terminology and the taxonomy and the house way of scoring a risk, baked into a small model at a precision that kept the facts.

02
It knows what was paged in for this decision.

The working set, retrieved for this tender and this buyer and this morning, and nothing larger.

03
It knows the state it is in.

Who is asking, under what authority, what the other agents have already done, and what the rules say it may not do.

04
And it knows what it has been told to forget.

Which is the part everybody skips, and the part that separates a system with judgement from a filing cabinet with a language model on top.

Yes, the filing cabinet is RAG. Retrieval-augmented generation is the second item on this list, done well, and it is necessary. But retrieval only ever adds. It has no opinion about what should stay out, no memory of what was deliberately excluded and why, and no way to prove that something was forgotten rather than merely not found. A system with judgement can show you the exclusion list. A filing cabinet cannot, because it never made one.

For thirty years we built software by encoding what computers should do. The next generation is about encoding what machines need to know at the moment they decide what to do. Which means the next generation of AI software will be judged less by how intelligent its model is than by how intelligently it decides what that model should know.

The advantage was never owning the biggest mind. It is building the system that knows what every mind needs to know, and when.

“Approve it.”

“Approve what?”

“Exactly.”

Thank you for reading.

Further reading

The earlier essays this one leans on, in the order they matter here.

That was the summary

Read the rest with your email

The full essay is about 30 minutes. Leave your email and the page unlocks here. Nothing is sent: the PDF edition, with contents, page numbers and a glossary, is a separate download at the end of the essay, yours when you ask for it.