The Paradox of the Empty Pocket
Why artificial intelligence is quietly a small business's best friend and a large enterprise's most expensive headache. MIT finds ~95% of corporate AI pilots show no bottom-line impact; Meta raises 2026 capex toward $145B, cuts 8,000 jobs, and admits agent progress 'hasn't accelerated the way we expected.' Meanwhile one seasoned operator ships 2,768 lines of tested code in a session. The capability isn't in question — the implementation is. On additive-vs-substitutive cost, ungoverned token spend, why the incumbent is circled by piranhas rather than one shark, and the governance gap a new category of agent-monitoring tools is being built to close.

Why artificial intelligence is quietly a small business’s best friend — and a large enterprise’s most expensive headache.
There is a very particular kind of vertigo that comes from watching an enormously rich organisation fail at something a single, deeply experienced operator can do with the right toolchain and the discipline to run it well. Not a novice with a laptop and beginner’s luck — someone who has spent thirty years moving through every rung of the ladder, from engineer to CTO to founder and most of the roles in between, now running Claude Code Max and a Codex subscription alongside tooling he built himself, on top of language models he has fine-tuned for his own coding work. That is not scrappy improvisation. It’s applied seniority, aimed at tools that didn’t exist for most of those thirty years. It shouldn’t be possible for that combination to outpace an organisation with unlimited capital behind it. It offends the natural order of things, in the same way it would offend you to watch a Formula 1 team lose a race to a driver who genuinely knows the circuit better than the pit wall does. And yet, in 2026, it is happening constantly, in full public view, in the quarterly filings of some of the most sophisticated technology companies on Earth. Something interesting is going on, and it is not, whatever the headlines suggest, simply that “AI doesn’t work.” It works extremely well. It just doesn’t work evenly, and the reasons why are far more revealing than the technology itself.
A disclosure before we begin, because this piece eventually lands on a problem I happen to build in: I make tooling in exactly this space, and one of my own products is an instance of the category of solution the argument arrives at. I’ve worked to keep the analysis standing on its own — the receipts below are real either way — but you should know the interest is there and weigh the conclusion with it in view.
The number that should worry every CFO
Let’s start with the sobering bit, because any good popular-science explainer earns the wonder by first being honest about the data. MIT’s Project NANDA looked hard at generative AI pilots across the corporate world and arrived at a genuinely startling figure: roughly ninety-five percent of them show no measurable impact on the bottom line whatsoever. Not modest impact. None. S&P Global Market Intelligence separately found the share of companies scrapping most of their AI initiatives had more than doubled in a year — from 17% to 42% — and PwC’s 29th Global CEO Survey found 56% of chief executives, more than half, reporting neither a revenue gain nor a cost saving from their AI spending at all. This is not a story about a handful of laggards. It is, depressingly, close to the median experience.
What makes this genuinely fascinating, rather than merely depressing, is why. The consistent finding across the research is that the failure is almost never about the model underperforming. It’s organisational: unclear success criteria, data that isn’t ready for the workflows AI is meant to slot into, pilots that never connect to an actual operating process, and — this is the one that should make every executive wince — a habit of measuring how many employees logged into the tool rather than whether anything got better because they did.
Meta, or: how to spend $145 billion and still not be sure it worked
If you want to see this paradox at true scale, there is no better case study currently running than Meta. The company’s 2026 capital-expenditure guidance began at $115–135 billion, overwhelmingly aimed at AI infrastructure, and was subsequently revised upward toward $145 billion; in the same season it announced it was cutting roughly eight thousand jobs — about a tenth of its workforce — with Mark Zuckerberg telling staff directly that compute and people are now the company’s two dominant cost centres, and one of them was winning. Meta’s first-quarter costs for 2026 jumped more than a third year on year, driven overwhelmingly by data centres and depreciation rather than payroll, and the company’s free cash flow is on course to collapse from over $43 billion to roughly $8 billion in a single year as the capital spending eats into everything else.

Here is the detail that really deserves to be sat with, though. In July 2026, at an internal town hall, Zuckerberg told employees that AI agent development over the preceding four months “hasn’t really accelerated in the way that we expected.” That is about as close as a chief executive of a trillion-dollar company gets to saying, out loud, in front of the people he’d just laid off in the technology’s name, that the thing hasn’t yet done what it was purchased to do. It is worth pausing on the sheer strangeness of that sentence existing in the public record at all.
And it isn’t only the infrastructure. Meta separately had to send an internal memo to thousands of employees after its own staff’s day-to-day AI token consumption — not the giant training runs, just people using AI tools at their desks — quietly climbed toward billions of dollars a year, with leadership openly warning against what the company nicknamed “tokenmaxxing”: people running up usage numbers on gamified leaderboards that looked like progress but weren’t. Uber told a strikingly similar story: it burned through its entire annual AI coding budget in four months and had to cap spending per employee, despite the overwhelming majority of its engineers using the tools monthly and a majority of new code being AI-generated. Its own leadership admitted, refreshingly honestly, that the link between all that token spending and any measurable output “is not there yet.”
Why the unencumbered operator doesn’t have this problem
Here’s where the physics of the thing becomes genuinely elegant, if you’re the sort of person who finds elegance in spreadsheets. A large enterprise adopting AI is not replacing a cost, in the way the marketing decks imply. It is almost always adding one, on top of a payroll it already carries, a legacy technology estate it already maintains, and a governance layer built to manage humans rather than autonomous software agents. Ninety-five thousand engineers do not become five thousand engineers plus some tokens the week the company buys a licence. The people are still there, still salaried, still requiring management, office space, health insurance and performance reviews — and now there is also a token bill on top, frequently an ungoverned one, because nobody built the metering and accountability structures fast enough to keep pace with how enthusiastically the tools got adopted. Meta didn’t lay off eight thousand people because AI made those roles obsolete overnight; by Zuckerberg’s own account, it was to help offset the infrastructure bill AI itself created. The maths only works if the AI eventually delivers enough acceleration to justify a leaner structure — and per Zuckerberg’s own July admission, that acceleration was, at the time he said it, not yet showing up.
A small business, or a single person building alone, faces none of that arithmetic. There is no incumbent headcount to protect or displace, because there was never enough money to hire the headcount in the first place. There’s no six-layer governance process to route a new tool through, because there’s one person deciding whether to use it, and they decide on a Tuesday afternoon and are using it by Wednesday. There’s no cultural resentment to manage, because nobody in the building is quietly wondering whether the software is coming for their job — the software is the extra colleague nobody could otherwise afford. When a solo founder builds a working product for the price of a few API calls, using a tool that took an afternoon to learn, they are not disrupting an existing cost structure. They are becoming one, from nothing, for the first time — which is precisely why the leverage looks so absurd next to a hyperscaler’s balance sheet. It isn’t that the small operator is smarter. It’s that they have no legacy to carry, no incumbent to protect, and nothing to lose by simply trying the thing on Tuesday afternoon.
A fair reader will push here, and rightly, so let me be precise about the claim, because the loose version is easy to knock down. The hero of this story is not “small business” in the abstract — most small businesses do not have a thirty-year ex-CTO at the keyboard, and the sheer leverage on display leans heavily on the operator, not merely the org chart. What actually generalises is narrower and far sturdier: the advantage of being unencumbered. A senior operator carrying no legacy will out-run a large one carrying all of it, because the tool collapses the cost of trying, not the need for judgement — the expertise is what turns that speed into shipped product, and the absence of legacy is what lets the speed exist at all. So read the rest of this piece as a claim about unencumbered operators who happen to know what they are doing, not about small size as a virtue in itself. That version is both truer and much harder to argue with.
The honest ledger from the other side of the desk
It would be too neat, and a little dishonest, to end the small-business argument on a note of pure triumph — as if the resource-less builder gets all the leverage and none of the bill. They don’t. Pull the receipts from four consecutive days of one person’s own agentic coding work — debugging, automated test runs, tidying up several codebases in parallel — and the sessions read: $105.84, $71.32, $13.26, $314.56, $24.16. Add them and you’re at just over five hundred dollars, for one person, in under a week, not a month. That is not nothing. For someone bootstrapping alone, it’s a genuinely felt number.

What’s revealing is why it stacks up, because the underlying causes are the same ones sitting inside Meta’s internal memo, just at a scale of one desk instead of ninety-five thousand. The tool’s own usage breakdown points at exactly the same behaviours: sessions run deep into long context windows, which cost more even when partially cached; a majority of the spend comes from subagent-heavy sessions, where each spawned subagent quietly runs its own separate bill; and a meaningful slice comes from running several sessions in parallel, all drawing against the same shared limit. Swap “subagent” for “employee experimenting with a new tool” and “parallel sessions” for “six teams running six uncoordinated pilots,” and you have, almost word for word, the exact pattern behind Meta’s “tokenmaxxing” memo and Uber’s four-month budget exhaustion. The physics of ungoverned AI spend doesn’t care whether the org chart has one name on it or ninety-five thousand. It’s the same curve, just measured in one person’s evening instead of one company’s quarter.
The difference that actually matters isn’t that the solo builder is somehow immune to the cost. It’s what that five hundred dollars buys, set against what it displaces. For an enterprise, that same dollar is additive — one more line on top of the payroll, the real estate, the pension contributions, and everything else that was already there before anyone touched an AI tool. For the person paying it out of pocket, it is the payroll. It’s the entire engineering department, the QA function, and the DevOps hire that could never otherwise be afforded, compressed into a single, closely-watched invoice. Both figures sting. Only one of them was, structurally, ever going to have somewhere cheaper to be.
What a month looks like, and what a floundering team does to that number
Stretch that four-day ledger out honestly and the picture sharpens further. Five hundred and twenty-nine dollars across four days works out to a little over $130 a day. Run that same pace for a full month, not because anyone would sustain peak-intensity debugging every single day, but as a ceiling worth having in view, and you land at just under $4,000 for one disciplined, technically fluent person running heavy agentic workflows across several codebases.
Now put that same daily rate inside a large engineering organisation, and change nothing else about the headcount. This is the scenario worth sitting with, because it’s the one actually playing out inside the companies posting the ROI statistics from earlier in this piece: nobody gets let go, the org chart stays exactly as it was, and a new capability simply gets handed to everyone with a login. Imagine a two-hundred-person engineering team, and imagine — generously, because this is roughly what “floundering” looks like in practice — that only a fifth of them are running sessions anywhere near this intensively, without the discipline to manage context length, subagent sprawl, or parallel sessions, simply because nobody told them how, or why it mattered. Forty developers at roughly $4,000 a month each is somewhere north of $150,000 a month, comfortably past $1.8 million a year — and to be clear, that is a modelled ceiling, not a measured invoice: an illustration of the shape of the problem rather than a bill anyone has actually paid. The point survives the arithmetic either way, because that spend buys the organisation precisely nothing in reduced headcount — the same two hundred salaries are still being paid underneath it. It is pure addition, not substitution, exactly the trap this whole piece has been circling.

This is, almost exactly, the shape of the problem Meta’s leadership flagged internally and the one Uber solved with a blunt per-employee spending cap: not that the tool is too expensive to be worth having, but that handing a genuinely powerful, genuinely uncapped capability to a large group of people who were never trained to use it with any discipline produces a bill that scales with enthusiasm rather than with output. A solo builder learns the cost discipline fast, because every dollar is felt personally and immediately. A two-hundred-person team, floundering gently in the background with the same people still in every seat, learns it only once someone in finance asks, months later, why the AI line item has quietly become larger than several people’s salaries combined — and even then, discovers that undoing the habit is far harder than avoiding it would have been.
Categorically proven, catastrophically implemented
Let’s be precise about what the evidence in this piece actually shows, because it would be easy to walk away from the Meta numbers thinking the technology itself is the problem. It isn’t. The proof that AI categorically works sits earlier in this very article, in numbers nobody had to model or estimate: seventeen working products built by one person over five years, a $90-and-mostly-CodeEasy build that unlocked an entire second venture, and four days of receipts — real ones, not projections — showing a single developer shipping 2,768 lines of working, tested code in one session alone. That isn’t a pilot. That isn’t a demo that impressed a steering committee and then quietly died. That’s output, delivered, on the record. The capability is not in question. What’s in question, almost entirely, is what organisations do with it once they’ve bought it.
And what most of them do with it is one of two equally unproductive things. The first is to take a genuinely capable tool and shoehorn it into something so narrowly scoped, so heavily filtered through procurement and compliance and legacy integration requirements, that what comes out the other end barely resembles the thing it started as. This isn’t about any particular vendor — it’s about the rollout, not the tool — but there is a very real pattern of what you might call assistant-grade deployments: a genuinely capable AI coding assistant rolled out with its capability quietly narrowed to something not far from autocomplete with a chat window attached, followed by a puzzled search for the productivity gains that never showed up. That kind of deployment doesn’t solve a business problem. It solves a perception problem — the org gets to say it “has AI” in the board deck — while the actual bottleneck underneath sits exactly where it was.
The second failure mode is the mirror image of the first: handing someone a genuinely powerful, largely unconstrained tool like Claude Desktop, purely because a security team somewhere decided that was the “safe” option, with no framework for what governed, monitored, accountable use of that capability actually looks like inside the organisation. Underpowered-and-shoehorned on one side, powerful-and-ungoverned on the other — and both are symptoms of the same underlying gap: nobody has built the layer that sits between “the model is capable” and “the organisation can actually trust and account for what it’s doing.” That is precisely the gap tools like SILO.RED exist to close — infrastructure built around a simple, uncomfortable question: right now, hundreds of thousands of AI agents are already operating inside enterprise networks, some with root access, reading code, databases, and internal communications, and almost nobody can currently answer who, if anyone, is watching them. Until that layer exists and is actually adopted, buying a more capable model doesn’t fix the implementation problem. It just raises the stakes of not having solved it.

That gap is not a minor operational inconvenience. It’s an existential one. Every quarter an incumbent spends stuck in pilot purgatory — bolting neutered tools onto old processes, or handing out powerful ones with no governance — is a quarter a smaller, unencumbered, genuinely AI-accelerated competitor spends closing the distance. Here I’m openly estimating rather than measuring — treat this as a considered guess, not a data point — but it is not unreasonable to say that a large, slow-moving business carries a realistic risk of losing twenty percent or more of its business to competitors a fraction of its size, simply because those competitors never had the incumbent’s legacy weight to carry in the first place. The rational response to that risk isn’t to retreat further into procurement caution. It’s the opposite: large organisations should be actively opening strategic partnerships with exactly the small, nimble, AI-accelerated companies they’re at risk of losing ground to — not to acquire and neutralise them, but because those companies have already solved, out of sheer necessity, the implementation problem the enterprise is still struggling to name.
What enterprises should actually be doing with this
None of the above is an argument for large organisations to sit this one out, wait for the tooling to mature elsewhere, or outsource the entire problem to a vendor’s roadmap. Quite the opposite. The proof points already sitting in this article — seventeen products from one person, a build that cost roughly ninety dollars and was ninety percent machine-written, a four-day ledger showing real, shipped, working code — aren’t curiosities. They’re a demonstration that credible, narrowly-scoped, genuinely useful internal tooling, built around agents doing the underlying work, is now well within reach of a team that decides to actually build it rather than merely license something and hope. Quick wins of this kind are not a research problem any more. They’re an execution problem, and execution is exactly what a properly resourced enterprise should be better at than a solo operator, not worse.
That said, it would be glib to pretend scale doesn’t add real friction, because it does. An architecture review board is not bureaucratic theatre for its own sake — it exists because a large enterprise genuinely has more to lose from a badly integrated system, more regulatory surface area, and more stakeholders whose legitimate concerns (security, legal, data residency, existing system compatibility) all deserve a seat at the table. None of that friction is imaginary, and none of it should be waved away. But friction is not the same thing as impossibility, and that’s the distinction that matters here. A credible internal tool, built quickly around the same agentic capability a solo operator uses, and then run through the appropriate review rather than being blocked from starting one at all, is entirely achievable at enterprise scale. The problem this article has been describing isn’t that enterprises can’t do this. It’s that most of them aren’t yet organising to try, and are instead spending their energy on either neutering the tools they buy or leaving the powerful ones ungoverned — anything, it seems, other than the harder, more valuable work of building something narrow, credible, and their own.
That is where the focus needs to go, because the alternative has a shape worth sitting with honestly. Picture a large, well-capitalised incumbent as something swimming in water that looks calm from the surface. It isn’t being circled by one shark, one dramatic, headline-grabbing disruptor it can see coming and prepare a defence against. It’s surrounded by piranhas — dozens of small, fast, AI-accelerated competitors, each one only capable of taking a modest bite out of a single narrow slice of the business. No individual bite is fatal, or even alarming, on its own. Which is exactly the danger: nothing about it triggers the kind of response a single visible threat would. And then, not especially suddenly, the incumbent looks down and finds there isn’t very much of the business left. Not because of one battle it lost, but because it never quite got around to noticing it was being eaten a piece at a time.

The honest caveat, because a fair-minded piece would insist on one
None of this means large enterprises are foolish to invest, or that the technology is somehow a toy that only works for the under-resourced. If an organisation has already got the people — fully staffed, fully salaried, fully embedded — then AI adoption genuinely can be close to pure addition to the cost base rather than substitution, at least until the tooling, the governance, and the workflows mature enough to let headcount actually shrink in proportion to what the software can now do. That maturity gap is real, it is where the ninety-five percent failure figure lives, and pretending otherwise would be dishonest. The technology is not the obstacle. The obstacle is that a hundred-thousand-person organisation cannot turn on a dime the way one seasoned operator, running a well-chosen toolchain with real discipline, can — and every one of the governance failures — pilot sprawl, ungoverned token spend, tools bought without the workflow redesign to actually use them — is a symptom of exactly that structural rigidity, not of the models being insufficiently clever.
And there is a genuine other side to this ledger, worth granting plainly rather than in a footnote. An incumbent is not simply a slow solo operator with more staff; it holds things a solo builder structurally cannot conjure — distribution, brand trust, regulatory licences, balance-sheet endurance, data accumulated over decades, and relationships that took twenty years to earn. Those are real moats, and they are precisely why, for the roughly five percent who do get the integration right, AI as genuine substitution rather than mere addition is worth every ounce of the friction: applied at that scale, across that distribution, a real efficiency gain compounds into something no unencumbered individual can come close to matching. So the argument here isn’t that small beats big. It’s that big currently forfeits its own advantage every quarter it spends neutering the tools it buys or leaving them ungoverned — and that the very scale which makes the friction real is exactly what makes solving it so valuable. The thesis survives the steelman; it arguably needs it.
The quieter thesis underneath all of it
If there’s a single idea worth taking from watching Meta spend $145 billion while one seasoned operator, running the right toolchain with real discipline, gets further on a narrower problem for a fraction of the price, it’s this: artificial intelligence doesn’t equalise opportunity by making everyone equally capable. It equalises it by making the cost of trying fall furthest for the people who had the least capital to begin with, however deep their actual expertise runs. For a company with ninety-five thousand salaries already committed, a new tool is one more variable in an already impossibly complex equation. For someone with thirty years of hard-won judgement, the right tools, and an evening free, it is the entire equation, solved in one sitting. That inversion — the very largest organisations on the planet struggling with the exact tool that is quietly transforming what one experienced, well-equipped person can build — is not a contradiction. It’s the whole story of this moment, told twice, at two wildly different scales, with the same technology and two entirely different starting conditions.
Sources
- MIT Project NANDA — The GenAI Divide: State of AI in Business 2025 — the finding that roughly 95% of enterprise generative-AI pilots show no measurable P&L impact. Report PDF · Fortune coverage
- S&P Global Market Intelligence — Voice of the Enterprise: AI & Machine Learning (2025) — the share of organisations abandoning most of their AI initiatives rising from 17% to 42% year on year. spglobal.com
- PwC — 29th Annual Global CEO Survey (Jan 2026) — 56% of CEOs reporting neither higher revenue nor lower cost from AI. pwc.com
- Meta — 2026 capital-expenditure guidance and quarterly results; Mark Zuckerberg’s internal town-hall remarks on AI-agent progress and on compute and people as the two dominant cost centres.
- Uber — reported exhaustion of its annual AI-coding budget in four months, and the subsequent per-employee spending cap.
- The four-day cost ledger, the seventeen-product portfolio, and the ~$90, roughly 90%-machine-written build are the author’s own primary records.