The Economics of Thought
The industrial revolution reduced the cost of labour. The computing revolution reduced the cost of calculation. The AI revolution appears to be reducing the cost of judgement — and I noticed the day a quota meter reset underneath me while I was doing the washing up. Two screenshots fourteen hours apart, a weekly counter going 100% to 0% with the reset date unmoved, a $1,497 day, and 1.3 billion cache reads against 2.4 million fresh tokens — which is the real finding, because it means reusable cognition is becoming an economic asset. On the boundary condition where machine judgement reliably fails, and what one founder plus agentic systems does to sixteen products' worth of work that used to need sixteen companies. The original conclusion — that wisdom becomes scarce — was wrong. Wisdom does not become scarce. It becomes decisive.

The industrial revolution reduced the cost of labour. The computing revolution reduced the cost of calculation. The AI revolution appears to be reducing the cost of judgement.
Earlier this year I wrote The Paradox of the Empty Pocket, and I thought I had found the story.
The argument was simple enough. For most of history, capability followed capital. If you wanted to build something meaningful you needed people, infrastructure, management, and money. Companies existed, in large part, because coordinating capability was expensive — to build a product required specialists, to launch a platform required departments, to compete at all required scale. Then the cost of production collapsed. A laptop and a large language model became sufficient to create things that had previously required a floor of an office building. A founder with an empty pocket could suddenly command capabilities that had belonged exclusively to organisations.
That felt like the whole insight at the time. Capability had come unbolted from capital, and everything downstream of that would have to rearrange itself.
I was wrong. Not about the observation — about its importance. The more interesting story arrived several months later, in the least dramatic way imaginable, when Claude told me I had run out.
The day the meter reset
One afternoon, mid-flow, Claude Code informed me I had exhausted my utilisation. Fable was gone. The model quietly downgraded. The engineering workflows I had been running all day simply stopped, the way a lift stops between floors — no crash, no error, just the sudden absence of a thing that had been there a second ago.
Several hours later, Fable came back. The counter appeared to restart. Usage began climbing again from zero: nine percent, thirteen, sixteen. The scheduled weekly reset date, meanwhile, had not moved.
The event itself was almost certainly mundane. A quota recalculation. A capacity rebalance. A billing reconciliation. A rollout bug. I want to be unambiguous about this, because it is the least interesting part of what follows and I have no privileged insight into any of it: I do not know what happened, and I am not going to pretend I do.
But it forced a question I had never previously thought to ask, in fifteen years of buying software and thirty of building it:
What does thinking actually cost?
Not pricing. Not subscriptions. Not tokens.
Thinking.
I had been running what amounted to a full engineering organisation out of a terminal window for months, and I had never once asked what the underlying resource was, or what it was worth, or what happens to the world when its price starts to fall. The meter interrupted me hard enough that I finally looked at the meter.
Evidence from the edge
Before the argument, the exhibit. Everything that follows rests on what I actually saw, so it is worth putting the receipts on the table first and letting you check my working.
The work in question was a compliance and testing campaign against AgencioPredict — a signal-driven deep-quant prediction and trading platform, roughly 947,000 lines across 4,100 files and 2,000 directories. Not a toy. The kind of codebase where a masked test failure is not an inconvenience but a regulatory problem.
And it is worth being honest about what I was doing, because it is almost comically undramatic: I was tidying up loose ends. Clearing the desk before starting something new. This was housekeeping — the compliance checks and test sweeps you run not because they are interesting but because you do not want to carry them into the next thing. The entire argument that follows fell out of a chore.
The interruption, in two frames. This is the part I could previously only describe, and it is the reason any of this occurred to me. Two snapshots of the same machine, fourteen hours apart, on 16 July 2026:
| 00:16 | 14:21 | Scheduled reset | |
|---|---|---|---|
| Current week (Fable) | 100% used | 0% used | Jul 19, 3am — unchanged |
| Current week (all models) | 67% used | 2% used | Jul 19, 3am — unchanged |
Fable exhausted at midnight. Fable at zero by the afternoon. Week-wide usage down from 67% to 2%. And the scheduled reset date reading Jul 19 at 3am in both frames — the counters moved, the calendar didn’t. That is the whole anomaly, and I still cannot tell you what caused it.
The session, in numbers. $1,497.56 of reported value. 22 hours 52 minutes of API time across a wall-clock day and six hours. 49,398 lines added, 6,901 removed. And then the figure that reframed everything for me — the cache:
claude-fable-5: 1.8m input · 2.9m output · 691.8m cache read · 16.1m cache write
claude-opus-4-8: 570.2k input · 2.1m output · 614.0m cache read · 8.4m cache write
claude-haiku-4-5: 555 input · 15 output · 0 cache read · 0 cache write
Roughly 1.3 billion cache reads, against 2.4 million tokens of fresh input. For every token the system read anew, it re-read something it already knew about five hundred times over.
The shape of the work. 99% of usage from subagent-heavy sessions. 96% from sessions running eight hours or longer. 82% at over 150k context. Half of it under a general-purpose subagent role. That is not a person prompting a chatbot. That is a standing organisation with a headcount.
The behaviour, which is the part that actually matters. Across those hours the system was not writing code so much as running a process: warming routes sequentially so a cold start wouldn’t poison a benchmark, restarting daemons that had wedged, managing resource contention between parallel test runs, noticing that a shell command had reported success while a pipeline quietly masked a non-zero exit code. A full review pass surfaced ten defects. Nine were real. One it refuted itself, before showing me.
And one structural detail worth pausing on, because it is doing more work than it appears to. The two roles are not the same vendor. Claude implements; Codex reviews and approves — with the roles swappable, and an escalation policy that retries failed work on a more capable engine. The adversarial verification at the heart of this essay is not a model marking its own homework. It is one vendor’s model trying to find fault with another’s, continuously, because the cost of that argument has fallen to roughly nothing.

The evidence board above summarises a separate, smaller session from the same July 2026 campaign — $225.40 against 126 million cache reads, and the same 66.3% cache-dominated shape at one-tenth the scale. Two notes in the interest of not misleading anyone: the year stamps in the artwork read 2025, which is wrong — all of this is July 2026. And the 800K line count it shows was accurate when drawn; the codebase has since passed 947K, which is itself the point.
None of this is code generation. All of it is engineering process management, and I asked for none of it specifically.
Price is not cost
The first thing you notice, once you start looking, is that the number on the screen is not measuring what you assume it is measuring.
Fifteen hundred dollars is an eye-watering figure to attach to one person’s day. Except look at what the day consisted of: reading files, running tests, warming routes, restarting services, inspecting logs, waiting for builds. Twenty-two hours of API time against twenty-nine hours of wall clock — meaning that for a third of it, the expensive thing was sitting idle while my hardware did the work.
The visible price was clearly not the real cost. It never is. The ticket price is not the fuel cost. The cloud invoice is not the hardware cost. The API value is not the infrastructure cost. Somewhere beneath every pricing model lies an entirely different economic reality, and pricing is a story the vendor tells about that reality for reasons that include cost but are never limited to it.
That gap matters here more than usual, because the gap is where the argument lives.
The strange case of 1.3 billion cache reads
Look again at that breakdown, and hold the ratio in your head: 1.3 billion tokens read from cache against 2.4 million read fresh. Better than 99% of the activity in that session was not thinking. It was remembering.
If all 1.3 billion of those tokens were recomputed from scratch, every time, the economics would be impossible — not expensive, impossible, in the way that a city where every resident drives a private helicopter is not expensive but impossible. The only explanation that survives contact with the arithmetic is reuse. Something, somewhere, is refusing to think the same thought twice.
I cannot tell you how that mechanism is implemented, and I’ll come back to the discipline of not-knowing shortly. But I can tell you what it implies, and the implication is the single most useful thing I took from the whole episode:
The future economics of AI is not inference. It is memory.
The winning systems will not be the systems that think fastest. They will be the systems that avoid thinking twice.
This inverts how almost everyone talks about the field. The public conversation is about raw capability — bigger models, faster tokens, higher benchmarks, more parameters. But once you’re operating at the edge, the binding constraint stops being how well the thing thinks and starts being how much of its previous thinking it can retain, index and cheaply re-enter. Context is not a feature. Context is the balance sheet.
Which is why my sessions run 150k+ tokens of context across three hundred files and fifty active modules, and why that is economically viable at all. The cache hit rate does the work. Memory is what makes multi-hour autonomy affordable, and multi-hour autonomy is what makes everything else in this essay possible.
The great cost collapses
Step back far enough and economic history reads as a sequence of collapsing costs.
The Industrial Revolution reduced the cost of labour. The Computing Revolution reduced the cost of calculation. The Internet reduced the cost of communication. The Cloud reduced the cost of infrastructure.
Every one of them follows the same shape. Something scarce becomes abundant. Something expensive becomes cheap. Something that had required an organisation becomes available to an individual. And each time, the thing everyone argues about — the technology — turns out to be less consequential than the thing nobody notices, which is what the collapse does to the structures built on top of the old scarcity.
The AI revolution appears to be following the same path. But there is a rung in this ladder that almost nobody names, and the cache statistics are what put it there:
Labour
↓
Calculation
↓
Communication
↓
Infrastructure
↓
Memory ← the collapse nobody talks about
↓
Judgement
Memory belongs in that sequence, and it belongs immediately before judgement, because it is the precondition for it. Cheap judgement is not a consequence of better reasoning. It is a consequence of cheap recall. The cost of retaining, indexing and re-entering a hundred and fifty thousand tokens of accumulated understanding had to collapse before continuous judgement could become affordable — and once it did, judgement fell out of it almost as a by-product.
This is the part I think is genuinely new, and I want to state it plainly rather than leave it as an implication: reusable cognition is becoming an economic asset. Not the model. Not the inference. The retained, re-enterable state of having already thought something through. That is a form of capital, it sits on no balance sheet, and it is currently being accumulated by whoever happens to be running the longest sessions.
So the resource being transformed is a strange one.
Not wisdom. Not consciousness. Not creativity. Judgement — the plain operational business of analysing, comparing, validating, prioritising, planning, testing and deciding. The unglamorous middle layer of cognition that constitutes most of what most professionals actually do all day.

As with the evidence board, the snapshot panel here is drawn from the earlier, smaller session — 126M cache reads against an 800K-line codebase. The argument is the same at either scale; the larger session simply makes the ratio harder to argue with.
Judgement is more expensive than we think
Here is the thing that took me a while to see, and that I now cannot unsee.
Most organisations are not, fundamentally, collections of workers. They are collections of judgement.
Run down the roles that exist primarily because deciding is costly: architects, reviewers, auditors, QA teams, security teams, managers, risk committees, governance boards. Every one of them is an answer to a question that could not previously be answered cheaply. Is this design sound? Is this code correct? Is this claim true? Is this risk acceptable? Should this ship?
Historically, those answers required people, because judgement was scarce and lived only inside skulls. And when a resource is scarce, institutions emerge to aggregate it. That is what a company substantially is: a machine for pooling judgement, ranking it, routing questions to the right pool, and carrying the overhead of all the coordination that requires. The hierarchy exists because review is expensive. The committee exists because review is expensive. The six-week approval cycle exists because review is expensive.
So ask the obvious follow-up. What happens to a structure built entirely on the expense of review, when review stops being expensive?
I accidentally built an engineering department
I didn’t reason my way to any of this. I noticed it, the way you notice your accent has changed.
A modern agentic review no longer looks like this:
Find bug
Fix bug
Done
It increasingly looks like this:
Find issue
↓
Verify issue
↓
Attempt to refute issue
↓
Apply fix
↓
Run validation
↓
Detect validation flaw
↓
Repair validation
↓
Re-run validation
↓
Report outcome
Read that second workflow again and notice what it actually is. That is not code generation. That is a quality management system. Steps two and three — verify, then actively attempt to refute your own finding — are the entire epistemological content of a peer review process. Steps six through eight are a test engineer noticing that the test itself was wrong and fixing the harness before trusting the result.
At one point the agent worked out that a shell command had falsely reported success, because a pipeline had masked the real exit code. The build said green. The build was lying. It caught it.
Nobody told it to check. Nobody asked it to challenge the result. It simply did, because the cost of the additional check had fallen below the expected cost of being wrong.
That’s the whole thesis in one incident. The workflow evolved self-verification not because anyone designed a self-verifying workflow, but because verification had become economically viable. When judgement is cheap enough, doubt becomes rational. And a system that doubts itself, cheaply, repeatedly, is doing something organisations spend millions of pounds and several layers of management trying to achieve.
The strange economics of self-doubt
I want to dwell on this, because I think it is the most underrated development of the decade and it hides inside a boring word.
Humans avoid re-evaluation. Not from arrogance — from economics. Re-examining a conclusion you have already reached consumes a scarce and metabolically expensive resource, and every honest engineer knows the small, shameful voice that says it’s probably fine at 6pm on a Friday. That voice is not a character flaw. It is a rational response to the cost of thinking.
Organisations avoid re-evaluation for the same reason at a larger scale. A second review costs a fortnight and a meeting and someone’s goodwill. So we ration it, and we build elaborate structures — sampling, risk-weighting, materiality thresholds — whose real purpose is to decide which things we can afford not to check.
Agentic systems increasingly perform re-evaluation because for them, the cost of additional judgement is approaching zero.
That is not a small engineering convenience. That may be one of the more important economic developments of the century, and it is arriving disguised as a feature in a developer tool. An entire discipline of institutional design exists to compensate for the fact that checking things is expensive. What survives when it isn’t?
Look again at the review pass from the evidence board: ten defects found, nine real, one refuted in flight before it ever reached me. I did not adjudicate ten findings. I adjudicated nine, all of which were worth my time. The adversarial pass had already done the filtering that would once have been my job — and the filtering, not the finding, was always the expensive part.
What I can observe, what I can infer, and what I cannot know
A necessary interruption, because arguments like this one usually go wrong here, and I would rather be dull than wrong.
There is a strict line between what I have observed, what I can reasonably infer, and what I am simply not in a position to know. Most commentary in this space blurs all three into a confident story about How It Works, and most of that commentary is fiction with citations.
What is directly observable, from screenshots, logs and lived experience: Fable utilisation was exhausted. A downgrade occurred. Fable later became available again. Fourteen hours apart, the same machine reported the weekly Fable counter at 100% and then at 0%, and week-wide usage at 67% and then 2%, while the scheduled reset date read Jul 19 in both. A single session reported roughly 1.3 billion cache reads against 2.4 million fresh input tokens. Multi-agent workflows consumed the majority of session activity. Long-context sessions dominated overall usage.
What can reasonably be inferred: that some form of quota accounting system exists. That some form of cache reuse mechanism exists. That there is likely a distinction between real-time quota enforcement and longer-term usage accounting. That long-context agentic workflows are economically dependent on reuse and memory. These are interpretations of behaviour, not confirmed implementation details.
What I cannot know, and will not guess at: how Anthropic hosts its models. Whether its infrastructure is owned, leased or hybrid. How its quota systems are implemented. How cache reuse works internally. How usage is reconciled across services or regions. What runs underneath any of it. Anything I said about those would be speculation dressed as analysis.
There is a fourth category, and it is about me rather than the system. I am almost certainly an outlier. Whatever the median Claude user looks like, it is not someone running 22-hour API days at 99% subagent-heavy against a 947,000-line quant platform, holding 150k of context for 82% of their usage and pushing the thing until the meter does something strange. Most usage is a question, an answer, and done. Mine is an architecture review, fanning out to multi-agent analysis, then validation, refactoring, testing, and deployment prep — for hours, unattended.
That cuts both ways, and I want both edges visible. It means my experience does not generalise: nothing here should be read as what these tools are like for a normal person doing normal work, and the costs I am quoting are not costs anyone else is likely to see. But it is also precisely why there is anything to report. Quota systems, accounting systems, cache infrastructure, context management, long-running orchestration — these only reveal their shape when you lean on them hard enough to find the edges. Ordinary use never touches the walls. The behaviours in this essay are visible to me because I have spent a year pushing Claude to the extreme, and the edge is the only place the economics become legible.
And here is why the distinction is worth the paragraph rather than a footnote: the argument does not depend on any of it. The thesis survives even if every assumption I might make about the implementation is wrong. This is not a description of anyone’s architecture. It is an examination of the economic behaviour that becomes visible when you operate at the edge of a system, and that behaviour is the same regardless of what’s under the floorboards. Memory matters. Reuse matters. Orchestration matters. The cost of judgement is falling.
You do not need to know how the engine works to notice the world going past the window.
From tools to cognitive infrastructure
The old model was simple, and we have used it for seventy years:
Human
↓
Tool
↓
Result
The emerging model looks more like this:
Human
↓
Cognitive Infrastructure
↓
Continuous Judgement
↓
Result
The difference is not power. It’s persistence. A tool is a thing you pick up, apply, and put down; its involvement begins and ends with your instruction. What I have been working alongside doesn’t behave like that. It behaves like an operational layer — a continuous source of analysis, validation and challenge that keeps running when I stop looking at it, that pushes back on things I did not ask it to examine, and that occasionally tells me I am wrong about something I was confident about an hour ago.
Somewhere between the architecture reviews, the adversarial verification loops, the autonomous test-harness repairs and the security passes running inside my terminal, it stopped feeling like software.
Not an employee. Not a tool. Something in between.
A form of cognitive infrastructure — and I choose the word deliberately, in the same sense that roads and electricity are infrastructure. You do not use electricity the way you use a hammer. You build on the assumption of it. Its abundance stops being a feature of your work and becomes a property of your world, and everything you design afterwards quietly assumes it is there.
That is the transition happening now, and the reason the framing matters is that infrastructure changes what is thinkable, not merely what is doable.
The stack of the agentic age
None of this is a standalone argument. In retrospect, it sits inside a body of work I have been circling for a couple of years without seeing the shape of it, and the shape is a stack.
Read from the bottom. Trust Is the New Infrastructure is the foundation — without trust, autonomous systems cannot scale, full stop. The Agentic Operating System is the execution layer, where goals, memory, tools and agents become an operational system capable of pursuing outcomes rather than merely executing commands. The Paradox of the Empty Pocket is the capability layer, where capability comes unbolted from capital.
And this essay is the economic layer, which follows from the ones beneath it with an almost irritating inevitability:
Once cognition becomes executable, judgement becomes measurable.
Once judgement becomes measurable, it becomes priced.
Once it becomes priced, it becomes an economic resource.
The question is no longer whether machines can think. That question is spent, and it was always the wrong question anyway — a philosopher’s question wearing an engineer’s coat. The question that matters now is what happens when thinking itself becomes infrastructure.
The new organisational unit
For centuries, the company has been the fundamental unit of economic activity. It exists because coordinating capability is expensive, and because judgement had to be pooled to be affordable.
If judgement becomes abundant, that logic doesn’t disappear — but it inverts. The organisation stops existing primarily to accumulate thinking and starts existing to direct it. The bottleneck moves from performing the analysis to deciding which analysis is worth performing.
You can see the shape of the change in the contrast: the traditional organisation aggregates judgement, specialises functions, treats review as expensive, uses hierarchy for control, and lives with slow feedback loops. The emerging one orchestrates judgement, generalises capability, treats review as abundant, uses networks for alignment, and runs on continuous feedback. Not because it is more enlightened. Because the underlying price changed, and structures are downstream of prices.
The Paradox of the Empty Pocket argued that capability no longer requires capital. This essay extends it one turn:
Judgement may no longer require organisations.
The engineer commands a dozen specialist agents. The founder commands an entire virtual department. The organisation increasingly becomes a coordinator of cognition rather than a producer of it.
That does not mean organisations vanish. Almost nothing ever vanishes. It means the constraint moves — and every time in history the constraint has moved, the people watching the old constraint have badly misread the decade that followed.
The founder amplification effect
That claim deserves better evidence than one anecdote about an exit code, and I am in the slightly awkward position of being the evidence. So let me put my own portfolio on the table, with the obvious caveat that I am hardly a disinterested witness.
One person, no engineering department, currently carries sixteen products.
AgencioPredict — the platform from the evidence section — is a signal-driven deep-quant prediction and trading system running to 947,000 lines. SignalFabric is live alongside it: 506 signal primitives across 16 quant categories and 40+ integrations, fusing markets, macro, sentiment and physics-informed models into calibrated forecasts — then acting on them, describing a strategy in plain English, backtesting it and trading it live behind an AI risk critic and a kill switch. That is a quant desk. Quant desks have historically been floors of buildings.
Silo is a five-layer security architecture running from silicon to cloud, watching for compromised or rogue agents: 1.8 million lines of code, ten papers running past 240 pages, six platforms. That is a security firm plus a research division. It is also the foundation layer of the stack above, built to be enforced rather than assumed — trust is the new infrastructure is not a metaphor if you have to ship it.
Bertha is 29 microservices on EKS, multi-tenant, three years in the making on a patent-filed stack, running the whole marketing lifecycle from CRM mining through multi-LLM creative to omnichannel placement. That is an agency.
CodeEasy is a visual command centre for AI-assisted development with an autonomous multi-agent build pipeline behind eight stage gates. That is a developer tools company — and, pointedly, it is the tooling that makes the other fifteen tractable. The Claude-implements / Codex-validates arrangement from the evidence section is CodeEasy’s doing; the amplification is partly a product of having built the amplifier.
Sidekick turns conversation into queryable, provenance-backed long-horizon memory, gated so nothing acts without permission. Which, given everything above about memory being the real economics, is not a coincidence.
Add posit., MiFamilias, and the rest of the portfolio. Add the essays. Add the research. Add The Possible. And add The Governor, which isn’t in the sixteen at all — it’s a run governor for autonomous agents that I keep purely as a working explainer of the agentic operating system, the execution layer this whole essay sits on top of. That it exists is its own small piece of evidence: when capability is this cheap, you can afford to build something real for no reason other than that it explains something.
Now do the honest arithmetic. Any one of those would conventionally require a company: a founder, a CTO, an engineering team, QA, security review, technical writing, DevOps, and the management layer to coordinate them. Sixteen of them would require sixteen companies, or one large one with sixteen business units and the entire apparatus of internal coordination that implies — which is to say several hundred people and a nine-figure balance sheet before a line of code exists.
The formula the whole essay has been circling, stated as plainly as I can:
1 founder
+ agentic systems
+ reusable cognition
= previously organisational capability
I want to be precise about what this does and does not prove, because the triumphant version of this argument is easy to knock over and I would rather make the sturdy one.
It does not prove that anyone with a laptop can do this. The amplification is multiplicative, not additive — it multiplies whatever judgement you bring, and thirty years of moving through engineer, architect, CTO and founder is what I am bringing to the multiplication. Hand the same toolchain to someone without the scar tissue and you do not get sixteen products; you get sixteen plausible-looking things that fail in ways nobody notices until they matter. The agent can verify that the code is correct. It cannot tell you the product was a bad idea.
Nor does it prove the output equals what those sixteen companies would have produced. It plainly doesn’t. What it demonstrates is narrower and more interesting: the coordination overhead — the meetings, the handoffs, the review queues, the status reporting, the entire apparatus that exists because judgement had to be pooled across people to be affordable — has largely gone. Not been optimised. Gone. It turns out a substantial fraction of what an organisation does is not the work; it is the cost of getting judgement from the head that has it to the decision that needs it.
That is the founder amplification effect, and it is the closest thing to a controlled experiment I can offer: same person, same judgement, same thirty years — the only variable is whether the cost of judgement is scarce or abundant. The delta is fifteen extra products.
The boundary condition
Which brings us to the objection I have been deferring, and it is the right objection.
A reader who has come this far has every reason to ask: are we sure this is judgement at all, and not just very sophisticated pattern matching wearing judgement’s clothes?
I don’t think that question can be answered as posed, because it is the wrong shape. “Can AI judge?” invites a yes or a no, and both are wrong. The useful question is narrower: what kinds of judgement can these systems perform, and where do they fail? The answer, from a year at the edge, is unusually clean — and the line is not where most people assume.
Where it is genuinely strong:
- Verification. Determining whether a claim, a test, or a result is actually true. This is the one that surprised me most, and it is not pattern matching — refuting your own finding requires holding the finding and the counter-evidence simultaneously and preferring the evidence.
- Orchestration. Sequencing work, managing dependencies and contention, deciding what has to happen before what.
- Process correction. Noticing that the method itself is broken — the masked exit code, the test that passes for the wrong reason — and repairing the instrument rather than trusting the reading.
Where it is genuinely weak:
- Purpose selection. Choosing which problem is worth solving. Given a goal it is relentless; asked which goal deserves the year, it has nothing, and its fluency at sounding like it has something is precisely the hazard.
- Value formation. Deciding what should be traded against what. It will optimise any objective function you hand it and cannot tell you whether the function is the right one, because that judgement is upstream of anything it can compute.
- Consequence ownership. Carrying the outcome. Not predicting it — bearing it.
Notice the pattern. The system is strong wherever there is a ground truth to converge on, and weak wherever the answer depends on what you want. Verification has a fact of the matter. Purpose does not.
So the honest formulation of the thesis is not “AI can judge.” It is: the operational half of judgement is collapsing in cost, and the evaluative half is not. Everything downstream of a settled goal is becoming abundant. Everything upstream of it is exactly as expensive as it has always been.
Which is a much more defensible claim, and it sets up the thing I got wrong.
Wisdom is not judgement
Here is where I disagree with myself.
When I first developed this thesis, I concluded it neatly. Labour, then calculation, then judgement — therefore wisdom becomes scarce. It had the right rhythm. It landed. Every audience nodded.
The more I sat with it, the less I liked it. Not because it was wrong, exactly. Because it was lazy — it treated wisdom as though it were simply judgement with more experience stapled on, the next rung of the same ladder. It isn’t. They are different in kind, not degree.
Judgement answers: What should happen next?
Wisdom answers: Why are we doing this at all?
An agent may eventually determine the safest implementation, the most efficient architecture, the highest-probability outcome. Genuinely, reliably, better than I can, and I say that as someone who has spent thirty years developing exactly that muscle. But wisdom asks a different family of questions entirely: Should we build this? Should we launch this? What are we sacrificing? What will we regret?
Those are not optimisation problems. They do not have a loss function. They are questions of values, and values are not a harder version of analysis — they are a different thing that analysis cannot reach from the inside, no matter how much of it you buy.
And there is a second reason, which I find more interesting and slightly uncomfortable. Much of what we call wisdom is compressed experience: the accumulation of successes, failures, trade-offs, responsibility and consequences. An agent may possess extraordinary judgement. It may evaluate possibilities faster than any human who has ever lived.
Yet it does not bear consequences.
It does not lie awake. It does not carry the redundancy it recommended. It does not meet, five years later, the person whose project it killed. Wisdom is deeply entangled with consequence — arguably it is the residue of consequence — and a system that cannot be harmed by being wrong is missing the ingredient that turns judgement into wisdom in the first place. That is not a criticism of the technology. It is a description of what the technology is.
So if I were to amend my own conclusion today, I would change one word. Not:
Wisdom becomes scarce.
But:
Wisdom becomes decisive.
Wisdom does not become scarce because machines cannot possess it. Wisdom becomes more important because everything adjacent to it has become abundant, and abundance has a way of throwing the remaining scarcity into hard relief. Judgement may become infrastructure. Analysis may become infrastructure. Verification may become infrastructure. But purpose remains stubbornly difficult to automate — and purpose is where wisdom lives.
Purpose becomes the constraint
For centuries the bottleneck was labour. Then information. Then infrastructure. Now judgement itself is becoming abundant, in a terminal window, in real time, priced at fractions of a penny per verified thought.
When that happens, the limiting factor becomes purpose.
The question is no longer: Can we analyse this? We can analyse anything. We can analyse everything. We can analyse it nine ways, have the system refute its own conclusion, repair the test that proved it, and hand you a report before lunch.
The question becomes: Is this the question worth analysing?
And nothing in the stack answers that. Not the model, not the orchestration layer, not the cache, not the 1.3 billion tokens of retained context. That one comes back to you — which is either the most liberating or the most frightening sentence in this essay, depending entirely on whether you have thought about what you actually want.
The least interesting thing that happened
The strange thing about the utilisation meter was not that it reset.
The strange thing was that it made me realise I had stopped thinking about the model as software. Somewhere between the architecture reviews and the adversarial verification loops and the autonomous test repairs, it had quietly become something else, and I hadn’t noticed because I had been too busy using it to look at it.
The paradox of the empty pocket was that capability no longer required capital. The economics of thought may be that judgement no longer requires organisations.
If that is true, then the quota system is the least interesting thing that happened. It was never the story. It was the catalyst — the accidental interruption that made the water visible to the fish.
And I think there is something fitting in where it happened. I was not doing anything visionary. I was clearing the desk — running the compliance checks and the test sweeps you do at the end of a thing, so you can start the next one clean. The most significant economic observation I have made in years arrived while I was doing the washing up. That is usually how it goes. The shifts that matter don’t announce themselves in the work you were proud of; they turn up in the chores, which is exactly where you’d expect infrastructure to first become visible — at the moment it stops being remarkable.
Because the real story is not the cost of using these systems. The real story is that, for the first time in history, we are beginning to discover the market price of thought itself. And thought, it turns out, is cheaper than we assumed, and more replicable than we hoped, and worth precisely as much as the purpose we point it at.
The future may not belong to those who can think the fastest.
It may belong to those who can decide what is worth thinking about.