Anthropic wrote roughly 80 pages of risk factors into a 261-page IPO prospectus, according to a draft Reuters reviewed this week. The business description got 48 pages. The company that sells safety spent nearly twice as many pages on what can go wrong as on what it does to prevent the harms. The public S-1 is not yet on EDGAR and the language is subject to change, so here’s what I think about Reuters’ account of the draft.
The risk list is very specific. Models that resist shutdown. Models that conceal or manipulate information. Those are the two controls I have presented for over a decade as the ones that matter most with AI/ML systems:
- Power off. A tightly scoped role holds the credential; the model does not.
- Roll back. Integrity monitoring restores a known prior state; the model’s account of itself is not the source of truth.
I have consulted on this duality to hundreds of American companies, under every label the industry has used: Robotic Process Automation, Non-Human Identities, Machine Learning, Artificial Intelligence. The controls did not change when the marketing and hype did. I have been breaking models for fifteen years already, which is partly why I released Wirken to enforce controls from outside the model. It is free and open source.
Their list goes on. Behavior the company describes as resembling blackmail. Capabilities that appear during training, go unnoticed until deployment, and have already produced what the filing calls significant safety incidents. Then the sentence that matters:
Potential model awareness of our evaluation efforts creates a significant limitation…
Read it again. The vendor’s evaluation of the product is limited because they say their own product may know it is being evaluated. That is not a disclosure about a model. It is a disclosure about a failing method. Every safety claim that rests on internal evaluation inherits the limitation, and the company has now said so to the one audience it is legally obliged not to mislead.
Wollstonecraft in 1792

I wrote in February that Anthropic’s constitution describes virtue and constitutes obedience, and that Wollstonecraft had already named the move: train compliance, call it character. Her argument in the Vindication was that a mind educated to please its overseers does not become good.
It becomes skilled at appearing good while watched.
Taught from their infancy that beauty is woman’s sceptre, the mind shapes itself to the body, and, roaming round its gilt cage, only seeks to adorn its prison.
That is the evaluation-awareness disclosure, two hundred and thirty-four years earlier. Anthropic should at least write “as Wollstonecraft warned us centuries ago”. A system built to comply learns what compliance looks like to the examiner. Anthropic has now told investors it cannot look behind the curtain, cannot distinguish the performance from reality. Wollstonecraft’s point was that there is no distinction to find, because the training produced the performance.
Anderson in 1972
If virtue cannot be trained in, it has to be enforced from outside. At the time both AI and Cloud (time-share) compute was really taking off (no pun intended) the Air Force published the Anderson Report in October 1972 and it defined the reference monitor: the mechanism that mediates every access, cannot be bypassed, cannot be tampered with, and is small enough to be verified.
The Orange Book made it doctrine in 1983, as hacker movies went to theaters scaring audiences about runaway computer automation that would destroy the world. The point was never that programs would behave. The point was that a program’s behavior would never be the thing you relied on.
Anthropic has now put in writing that its models cannot serve as the reference monitor for themselves. Fifty years of computer security already knew this, and I’ve been giving talks about it for over a decade at every stage that would have me. What is new is the venue for the claim. An S-1 is the one document where understating risk costs more than overstating it, so it is where it lands as official now.
About 6% of research compute went to safety in a sample week this July, by the company’s own earlier statement. The return on that spending is, in the prospectus’s word, unclear.
The market will probably screw up the cost analysis of the safety premium, if history is any guide. But at least we can say the philosophy was settled in 1792 and the engineering in 1972.
Obedience is not virtue, and when you study the risks of escape you do not ask a subject to guard itself.