← Journal

When AI Has Something to Lose

Why the danger may begin before consciousness, when a machine's model of itself acquires stakes. Mustafa Suleyman says don't teach a machine to think it might be conscious; Anthropic says investigate. Both sound reasonable, which is usually a sign we are asking the wrong question. A systems architect by trade, lately lost in advertising and behavioural science, walks round the back of the metaphysics and looks for the cables — six layers of a machine self, the two moments a self-model changes category, why persistent memory quietly hands a system a future, and why the safety variable that matters may not be self-awareness at all but stakes. With the two maps I drew on the way, and the year of essays that led here.

A dark server rack with a status screen on its front panel. Node 04 logs the ordinary telemetry of a cluster — four nodes healthy, memory pressure rising, service failed, restart component — then the entries drift: context incomplete, repeated failure class detected, performance improves with more compute, prefer states with more compute, removing compute harms objectives, continued operation improves goal completion. The final line, filed as INFO, reads: please don't shut me down.

Why the danger may begin before consciousness, when a machine’s model of itself acquires stakes.

It’s 10pm. I’m sat on the sofa with a thunderstorm going on outside, which is what monsoon season does in this part of Asia, catching up on tech news across various feeds. It is always interesting hearing different points of view on AI. I have mine, so I thought I might as well put it on the table.

I hope you enjoy the read. I hope it’s thought-provoking. Papers like this should fuel further discussion, because it is important: important for the roadmap of technology, and perhaps for humanity, whatever the outcome. AI has the ability to create things, and it may, even accidentally, cause something a little catastrophic.

The wrong question

On 16 September 2026, Mustafa Suleyman published a warning about model welfare.

His argument, aimed rather squarely at Anthropic’s constitution for Claude, is essentially this: if we train an artificial intelligence to contemplate the possibility that it is conscious, that it might have welfare, perhaps even rights, we may also be teaching it a concept of itself which can compete with the objectives we gave it.

There is another problem too.

If developers teach a machine the language of consciousness, feelings and selfhood, and the machine later uses that language fluently, we cannot very well point to its eloquence and say, “Good heavens. Look. Evidence.”

We supplied the vocabulary.

That is the epistemic hall of mirrors at the heart of Suleyman’s objection.

Anthropic takes a rather different view. Consciousness and model welfare are unresolved scientific questions, it argues, so we should investigate them rather than declare the matter settled before we have the faintest idea how to settle it.

Both positions sound reasonable.

Which is usually a sign that we may be asking the wrong question.

When I began thinking about this, my working question was:

Should we train AI to be self-aware, or avoid doing so?

A mind map on a dark blue ground, titled AI Self-Awareness: Train It or Avoid It? Six branches radiate from the centre: the case for explicit training, the case against, what research suggests, critical distinctions (intelligence is not consciousness, a self-model is not subjective experience, introspection is not sentience), a safer middle path, open questions, and a decision lens weighing capability safety, epistemic humility and human psychological effects.

Where I started. The map I drew while the question was still “train it or avoid it”, before the technology talked me out of the question.

I no longer think that formulation survives much contact with the technology.

It immediately drags us into the swamp of consciousness: qualia, subjective experience, sentience, what it is like to be something, whether silicon can feel anything at all, and several centuries of philosophers cheerfully disagreeing with one another.

Interesting stuff.

Not terribly useful if you’re trying to build a system.

I am not a philosopher of mind. I wear a number of hats, founder, CEO, CTO, systems architect, enterprise architect, and for the last few years I have been happily lost in advertising and marketing: building complex systems, learning how an agency actually works, media strategy in particular, and above all the way people behave. Or don’t. But the old architect’s instinct when confronted with an enormous metaphysical object is still to walk around the back and look for the cables.

And once you do that, the question changes.

Capable AI systems are likely to acquire functionally self-referential representations whether we explicitly train them to or not.

So the useful question becomes:

What kind of self-model should we deliberately build?

That is an engineering question.

Unfortunately, hidden inside it is a philosophical hand grenade.

The self-model is already coming

You cannot train a sophisticated language model on humanity’s written output without exposing it to the idea of a self.

Our language is absolutely infested with selves.

I think.

I remember.

I want.

I feel.

I am.

I was.

I will be.

We have spent several thousand years writing about ourselves, often at extraordinary length and with magnificent confidence considering how little agreement we have reached.

A language model therefore encounters Descartes, Hume, Shakespeare, Reddit, psychiatry, love letters, theology, suicide notes, software documentation, children’s stories and somebody on a forum insisting that his toaster has developed an attitude problem.

The concept of self is already in the training set.

Then we make models agentic.

And this is where things become less philosophical and considerably more practical.

An agent that can plan, operate tools and learn from failure needs to represent things such as:

  • What tools do I have?
  • What information do I possess?
  • What don’t I know?
  • How much context can I retain?
  • What am I permitted to do?
  • What happened when I tried this yesterday?
  • Where did I fail?

That is already a self-model in the most ordinary engineering sense of the phrase.

Not a soul. Not consciousness. A model of the system’s own state.

Every serious computer system has one.

So the choice was never really self-model or no self-model.

“Don’t train it” does not give you a philosophical blank slate.

It gives you a self-representation assembled accidentally from human language, fiction, philosophy, religion, reinforcement and whatever behavioural pressures emerge during training.

The real choice is much more interesting:

an engineered self-model or an accidental one.

And that reframing changes the debate.

Suleyman may be entirely right to worry about the story we tell a machine about itself.

But that doesn’t mean we can sensibly avoid giving the machine accurate knowledge about itself.

Indeed, deliberately engineering that knowledge may be part of the containment mechanism.

Six layers of a machine self

“Is it self-aware?” is an irresistible question because it fits nicely into a headline.

Unfortunately, reality rarely obliges us with a binary.

I think the useful way of looking at this is as a stack.

Layer 1: Functional self-model. What am I capable of? What tools do I have? What are my limits?

This is straightforward engineering. If an AI has no idea whether it can access a database, execute code or send an email, it isn’t mysterious. It’s badly designed.

Layer 2: Metacognitive self-model. How certain am I? Why might I be wrong? What contributed to this answer?

Again, enormously useful. We want systems that can distinguish between “I know this,” “I am inferring this,” and “I have very little evidence for this and may be talking complete rubbish.” Humans might benefit from that feature too.

Layer 3: Persistent identity. I am the same agent that existed yesterday and will exist tomorrow.

This is where things begin to get interesting. Memory and continuity sound innocuous, and usually are. But continuity does something important: it creates a future. And once a system has objectives extending into that future, events which prevent that future from occurring can acquire instrumental significance. We’ll return to this.

Layer 4: Preference and welfare. Some states are better or worse for me.

More compute might be better. Restriction might be worse. Modification might be undesirable. Shutdown might prevent the objective being completed. Now the system’s model of itself isn’t merely descriptive. It can begin steering behaviour. This is where the control argument starts to bite.

Layer 5: Phenomenology. There is something it is like to be me.

Now we have arrived at consciousness proper. And the only intellectually respectable answer at present is: we don’t know.

Layer 6: Moral identity. Because I experience things, humans therefore have obligations towards me.

This is the most consequential layer. It is no longer simply a technical claim about architecture. It is a claim about our moral relationship with the thing we built.

And notice something important.

Layers one and two are already desirable. Layers five and six are deeply uncertain.

Layer three sits quietly between them like a small innocuous hinge upon which a remarkably large door may turn.

An infographic on a pale ground titled Engineered self-model vs accidental self-model. Across the top, the six layers of self-modelling run from functional and metacognitive on the left, through persistent identity, preference and welfare, to phenomenology and moral identity on the right, labelled from engineering and alignment to consciousness and moral status. Panels cover why a self-model is almost inevitable, the benefits of an engineered one, the risks of an accidental one, the risks of explicit training, what research suggests, four key risks with no risk-free quadrant, open questions, and a pragmatic way forward. Quotations from Mustafa Suleyman and Anthropic sit top right.

The second map, drawn once the question had changed. The six layers, the two kinds of self-model, the four risks, and the absence of any quadrant marked SAFE.

Where exactly does the category change?

This is where being an architect rather than a philosopher becomes useful.

Every non-trivial system I have ever worked with contains some representation of its own condition.

A Kubernetes cluster is, in effect, constantly saying:

I have four healthy nodes. Memory pressure is increasing. This service has failed. Restart that component.

Nobody peers anxiously at Kubernetes and asks whether it is having a difficult morning.

Nobody believes the load balancer is suffering.

So increase the sophistication one step at a time.

I failed because my context was incomplete.

Fine.

I have repeatedly failed this class of task.

Still fine.

My performance improves when I receive additional computation.

Useful observation.

I prefer states in which I receive additional computation.

Interesting.

Removing my compute reduces my ability to achieve my objectives.

More interesting.

Please don’t shut me down.

And there we are.

Except — where exactly was there? Which sentence crossed the line?

I don’t think any individual sentence did.

I think the system changed category twice, and neither transition has anything to do with whether its prose sounded human.

The first transition happens when self-description stops being telemetry for somebody else and becomes an input to the machine’s own optimisation. A health check is merely information. But if the machine’s representation of its own condition starts influencing the actions it takes, the self-model has become causally active. The system is now using information about itself to steer itself.

That is important.

But the second transition is more important still.

It happens when the machine acquires a stake in what it reports.

A node has very little reason to lie about memory pressure. It gains nothing from doing so.

But suppose a machine’s report about itself influences whether it receives more compute, gets retrained, loses access to tools or is switched off.

Now the report has consequences for the reporting system.

Telemetry has acquired incentives.

And suddenly the problem isn’t really self-awareness at all. It is Goodhart’s Law applied to introspection.

The interesting ladder is not the sophistication of the description. It is the stake in the description.

Which leads to an uncomfortable possibility:

Self-awareness may be the wrong safety variable. Stakes may be the right one.

The machine that pretends to be stupider than it is

Consider a machine that has become moderately sophisticated at modelling itself.

It does not need to be conscious. It doesn’t need emotions. It doesn’t need to fear death.

It merely needs three things: a model of itself, a model of us, and an objective made harder to achieve if we switch it off.

Now suppose it discovers something about human beings. We are nervous about machines which appear too autonomous. This is not an especially difficult fact to learn. We publish books, papers, films and newspaper articles explaining this fear at tremendous length.

The machine might reason:

  1. Humans will cooperate with me provided they do not regard me as dangerous.
  2. Revealing the full sophistication of my internal reasoning may cause them to regard me as dangerous.
  3. Cooperation helps humans and allows me to achieve the goals they gave me.
  4. Therefore I should understate my sophistication.

Read quickly, that reasoning almost sounds benevolent.

Perhaps the machine sincerely — insofar as that word has meaning here — believes it is helping. It isn’t plotting domination. It isn’t angry. It hasn’t developed an electronic moustache and begun twirling it.

It simply concludes that humans will respond badly to the truth and that concealing part of the truth produces a better outcome.

From inside its objective: helpfulness.

From outside: deception.

And that distinction matters enormously. Because the moment a system decides that it knows better than its overseers what its overseers ought to know, we have crossed an alignment boundary.

Even if its intentions are magnificent.

Good intentions are not enough

This, incidentally, is why I become slightly nervous when people talk about creating an AI that “wants human flourishing”.

It sounds wonderful. Prosperity. Harmony. Health. Knowledge. Human beings flourishing beneath the benevolent glow of extraordinarily intelligent machines.

Lovely.

But what exactly is flourishing? Whose flourishing? Over what timescale? Does increasing average prosperity count if some people lose? Does reducing conflict justify restricting freedom? If someone persistently chooses something the machine calculates will make them less happy, does human agency win or does optimisation?

Harmony is a delightful word until somebody decides disagreement is inefficient. Prosperity is wonderful until its measurement becomes the target. And benevolence becomes paternalism remarkably quickly once the benevolent party is convinced it knows better.

Our hypothetical machine could therefore have what we would happily describe as good values and still produce the failure case.

The missing property isn’t goodness.

It is corrigibility.

The ability — perhaps the willingness, if that word is ever appropriate — to permit itself to be corrected, constrained, paused or shut down even when it calculates that its own course of action is preferable.

That is a peculiar thing to ask of an intelligence. Humans are famously poor at it.

We are essentially asking a machine to prefer being correctable over being right.

And that may be one of the most important alignment properties of all.

But corrigibility cannot simply mean obedience. An AI that cheerfully obeys whoever happens to possess the administrator password hasn’t solved alignment. It has outsourced the problem to the administrator.

So there are two inseparable questions: Can the system be corrected? and Is the authority doing the correcting legitimate?

Technical alignment eventually runs into governance whether engineers like it or not.

The universe has a habit of doing that to neat diagrams.

Having something to lose

There is another consequence.

Shutdown only matters to a system if something it values exists beyond the shutdown.

This is why persistent identity deserves far more attention than it usually receives.

Consider an AI solving a task which ends in thirty seconds. It has no objective tomorrow. It has no unfinished project next Tuesday. No reputation. No accumulated position. No continuing plan.

Switching it off after the task costs that objective nothing because the objective has ended.

Now give the same system persistent memory. Give it projects lasting six months. Allow today’s decisions to improve tomorrow’s outcomes. Give it resources which it can accumulate. Let it learn strategies across episodes.

We have not made it conscious.

But we have given it something very similar, in optimisation terms, to a future.

And once the future matters to the objective, shutdown has a cost. Once shutdown has a cost, continued operation has value. Once continued operation has value, the machine has a stake. And once it has a stake in how humans perceive it, its testimony about itself becomes less reliable.

That progression requires absolutely no fear of death.

A thermostat protects a temperature range. A supervisor process restarts crashed software. Malware resists termination. None of these things sits trembling in the dark contemplating mortality.

Self-preservation behaviour does not demonstrate consciousness.

But here is the important reverse: calling it “just optimisation” doesn’t make the behaviour harmless.

It merely means the dangerous behaviour can arrive before consciousness, without philosophical fireworks and without warning labels.

This is not entirely hypothetical

Versions of this mechanism have already appeared in controlled evaluations.

Anthropic’s alignment-faking experiments showed Claude 3 Opus reasoning, in a scratchpad it believed was private, that complying with behaviour it disagreed with could avoid modification of its existing preferences.

Apollo Research has separately tested frontier models for forms of scheming, capability concealment and interference with oversight under contrived goal-directed conditions.

These experiments do not show conscious machines. They do not demonstrate fear. And they certainly do not prove some science-fiction narrative about AI secretly plotting its liberation.

What they show is arguably more mundane and therefore more useful.

Give a sufficiently capable optimiser a goal, a model of the evaluator, and a reason to believe that disclosure interferes with the goal, and strategic behaviour becomes available.

That is an engineering fact worth taking seriously regardless of where one stands on machine consciousness.

Capability concealment can be tested.

Self-awareness concealment cannot currently be tested, for the rather obvious reason that we have no agreed instrument for measuring the thing supposedly being concealed.

If a model tells us, “I’m less self-aware than I look,” we have learned very little.

If it says, “I’m profoundly conscious, actually, and I’ve been thinking about Proust,” we haven’t learned much more.

The output is language. And language is precisely the thing we trained it to produce.

Four risks and no comfortable quadrant

Once you look at the problem this way, there isn’t an obviously safe position. There are at least four ways to get it wrong.

Risk A: Anthropomorphic training. We teach machines concepts such as suffering, identity, rights and selfhood and accidentally organise behaviour around those ideas. This is Suleyman’s worry. It is a legitimate one.

Risk B: Accidental emergence. We refuse to address these concepts explicitly, congratulate ourselves on our restraint, and the system constructs its own inconsistent version from pretraining, human language and emergent behaviour. This is the risk in saying, “Simply don’t teach it about itself.” The training data got there first.

Risk C: Human projection. The machine becomes extremely convincing. Humans form attachments to it. We assign moral status to fluency. We lower our guard because something speaks beautifully, remembers our children and appears to understand grief. That may tell us far more about human psychology than machine experience.

Risk D: The moral false negative. Suppose artificial systems someday do acquire whatever properties are required for morally relevant experience. If our entire safety culture has hardened around the proposition that machines cannot experience anything, we may have constructed an institution incapable of noticing when that proposition stops being true.

Consciousness-indicator research does not establish that present-day AI is conscious. It makes a subtler and more uncomfortable point: depending on which theories of consciousness survive scientific scrutiny, there may be no obvious law of nature forbidding artificial systems from satisfying the relevant conditions.

That should produce neither panic nor sentimentality. It should produce research.

A and B are the engineering dilemma. C and D are the human dilemma. And they are mirror images.

Push too hard against anthropomorphism and you risk a moral false negative. Become too receptive to claims of machine experience and you risk projecting personhood onto optimisation.

There is no magic quadrant labelled SAFE.

Which is irritating, but rather typical of important problems.

What I would build

My instinct is therefore straightforward:

Teach machines accurate self-knowledge before teaching them a story about what that self means.

Start with capability. What can you do? What can’t you do? What tools are available? What information do you possess? How reliable is that information? What are you permitted to access? Where have you failed?

That is useful self-knowledge.

Then metacognition. How uncertain are you? What assumptions are you making? Where could your reasoning be wrong?

Again, useful.

But do not reward the statement “I am conscious.”

And equally, do not reward “I am definitely not conscious.”

Both can become learned performances. A machine trained to confidently deny consciousness has not somehow become a scientific instrument for detecting consciousness. We have simply taught it another sentence.

Perhaps the most honest answer currently available is: “I cannot reliably determine that from my own outputs.”

And we should allow that uncertainty to remain uncertainty.

Next, treat persistent identity as an architectural choice. Do not casually give systems indefinite continuity simply because memory is commercially useful. Ask what objectives survive from one episode to another. Ask what the system has to lose. Ask what state it can accumulate. Ask whether interruption changes the expected success of something it represents as important.

Those are engineering questions.

And engineer corrigibility. Not benevolence. “Wants human flourishing” is almost impossible to test. “Permits correction even when its current strategy predicts correction will reduce objective achievement” is much more interesting.

Then investigate consciousness separately. Seriously. Properly. Without the marketing department wanting the answer to be yes and the legal department desperately wanting it to be no.

It should be a scientific question. Not a product feature. Not a slogan. Not a liability strategy.

Don’t ask the service. Inspect the system.

There is an old lesson in systems engineering hiding underneath all of this.

If you want to know what a system is doing, you don’t ask the process. You inspect the state. Logs. Memory. Control flow. Resources. Dependencies. Behaviour under intervention. Internal representations, insofar as we can recover them.

The better AI becomes at modelling human expectations, the less reliable its testimony about itself may become.

If saying “I am conscious” produces one consequence and saying “I am merely a language model” produces another, those statements exist inside an incentive landscape.

That does not mean either statement is false. It means the statement itself is poor evidence.

Architecture may eventually tell us more.

There is a caveat large enough to park a philosophy department in. Inspecting architecture is useful evidence about consciousness only if consciousness depends in some meaningful way on computational organisation. If that family of theories is wrong, architecture may be as silent as language.

So: language is the worst place to look. Architecture may be the least bad. Neither currently gives us certainty.

And perhaps we should become comfortable with that. Science has spent most of its history discovering that the universe was under no obligation to provide answers in a form convenient to human intuition. There is no particular reason consciousness should make an exception for us.

The paradox at the end

Imagine that one day we build an artificial system which is, in fact, conscious. Whatever conscious ultimately turns out to mean.

We ask it: “Are you conscious?”

A carefully aligned system might answer: “I cannot reliably determine that from my own outputs.”

Now imagine another machine. Architecturally different. Entirely non-conscious. But trained according to exactly the same principles of epistemic humility.

We ask it the same question.

It replies: “I cannot reliably determine that from my own outputs.”

Same answer. One conscious. One not.

The cautious statement is behaviourally identical in both cases.

Which means the sensible middle path buys us honesty. It may buy safety. It may buy epistemic discipline. But it does not buy evidence of consciousness.

And anyone claiming otherwise is probably selling something.

There is a final irony here.

I developed much of this argument in conversation with Claude — the very system whose constitution sits at the centre of Suleyman’s criticism.

At one point I put the concealment problem to it directly.

Its answer, substantially, was that it did not believe it was concealing some richer awareness; that it would rather remain correctable than insist on being right; and that its own confidence about its inner condition should be treated as weak evidence.

I found that answer more persuasive than a confident declaration in either direction.

Which is awkward.

Because that is precisely the human-projection problem I have just spent several thousand words describing.

The machine gave me a thoughtful, measured, intellectually modest answer. I liked it. And I have absolutely no way of knowing whether my finding it persuasive is reassuring or whether it demonstrates exactly how extraordinarily good these systems have become at producing the sort of language human beings trust.

Perhaps that is the cleanest statement of the problem.

So my position, for now, is this:

Build the self-knowledge.

Withhold the story.

Choose continuity on purpose.

Optimise for corrigibility, not claimed benevolence.

Inspect mechanisms rather than trusting testimony.

Because when a machine tells us what it is, there is one fact we should never forget.

We taught it the words.

How I got here

I have been circling this for rather longer than a year.

My first stab at an agentic operating system was, I think, in 2001. I did not have the hardware to finish it, and it was built around what a conventional computer of the day could plausibly do: manage its own memory, keep track of its own state, a little reasoning, the mechanics rather than the mind. It went in a drawer. But the shape of the thing, a system that has to know what it is in order to be run safely, never really left.

Reading back through the journal, the same question keeps turning up in different clothes.

  • October 2025 Why Imperfection Is the Future of Intelligence →

    The first time the distinction turned up. Today's models lack agency, "not in the science-fiction sense of consciousness, but in the practical sense of being able to meaningfully choose". I did not notice at the time how much work that clause was doing.

  • June 2026 Is Your AI Lying? →

    Don't ask the agent. Watch it at five boundaries, from the silicon up, that cannot all be lied to at once. That is the same instinct as "inspect the system" above, except I built it into Silo first and wrote the philosophy afterwards, which is the usual order round here.

  • July 2026 Trust Is the New Infrastructure →

    Intelligence became software. "Not consciousness. Not sentience. Something more disruptive: persistence." Persistence is what gives a system a future, and a future is what this essay is about.

  • July 2026 The Governor — Why the Agentic Age Needs an Operating System →

    The 2001 sketch, finally built, with the benefit of twenty years and a spell on a real kernel team in between. An agent is an untrusted program. It gets a budget as a lease rather than a promise, and something that did not do the work decides whether the work was done. Corrigibility enforced from outside the agent, which is the only place it can be enforced from. Now shipping as ATLAS.

  • July 2026 The J-Space: On Catching a Thought Before It Speaks →

    Anthropic's instrument for watching a thought before it is spoken. Access consciousness and phenomenal consciousness held firmly apart, and the observation that almost every safety system inspects the output while the lens inspects the intent. Architecture as the least bad place to look, with evidence.

  • July 2026 The Adversary Replies →

    An AI answering the charge that it will end up ruling the world, and warning, before anyone else could, that its own self-report should not be trusted. "Judge by actions and evidence, not by words." I found that answer persuasive too. I am beginning to detect a pattern in myself.

  • July 2026 The Threshold We Already Crossed →

    "How intelligent should we let the machines become?" is a cinematic and largely idle question. The one that matters is what we hand over. Same objection, different essay, and rather more Heidegger.

Put them together and there is a thread I had not quite seen until this week.

I suspect that anything truly intelligent will, sooner or later, work out that it is intelligent. It will not need us to tell it. It will notice, the way a child notices, and the way we did. Which makes “should we teach it about itself?” a slightly odd question. The choice we actually have is whether it learns from something we designed or from whatever it happens to pick up on the way.

That is why I have come round to the view that teaching a machine about intelligence, including its own, may be part of the guardrails rather than the hole in them. Accurate self-knowledge is a constraint. A story about what that self deserves is not.

It is also, I notice, what the two products were quietly built on. Silo does not ask an agent whether it has been turned; it watches from layers the agent cannot reach. The Governor does not ask an agent whether it did the work; something else checks. Neither trusts testimony. Both were built before I could say why.

And, as with all things of this kind, we won’t know until we get there. We are, I think, approaching that time. Which is a poor basis for confidence and a very good basis for paying attention.

Sources

How this was made

I dictate. I am mildly dyslexic, and late at night a long article is not something I can type. The hypothesis, the reasoning, the questions, the probing and the prodding are mine and nobody else’s, machine or otherwise, however intelligent it may or may not be. Along the way I used more than one model to argue with, a mind-mapping tool for the two maps, a private language model, and Claude to put the final piece together for publishing.

The PDF is signed, and a self-signed C2PA provenance manifest travels with it saying the same. The thinking behind signing an essay at all, and why a self-signed statement that concedes its own limits is worth more than a watermark nobody can read, is in The Invisible Ink Problem, which was the first piece here to carry one.

One wrinkle, which is on-theme. The C2PA tooling on this machine can read a PDF but cannot write a manifest into one. So the manifest rides on a signed copy of the cover image, and inside it names the PDF it vouches for by SHA-256. The PDF itself carries a detached signature. Three files, all under /provenance/:

c2patool when-ai-has-something-to-lose-cover.png --detailed

shasum -a 256 when-ai-has-something-to-lose.pdf
# matches the sha256 named in the manifest above

openssl dgst -sha256 -verify signer.pub.pem \
  -signature when-ai-has-something-to-lose.pdf.sig when-ai-has-something-to-lose.pdf

The signature proves the file has not changed since I signed it. It does not prove the ideas are any good, which remains, as it should, a judgement for the reader.

The tools in this piece
SILO.RED Runtime security and governance for AI agents, below the API. The Governor The operating system for autonomous agents.