All posts by Davi Ottenheimer

NVIDIA Restates Wirken as Five Principles, Then Sells the Sixth as a DPU

NVIDIA opened its agent safety announcement with the browser. The web became safe, it says, when the browser stopped trusting the page. That is the right history.

Then they try to sell you the reverse of it.

Not so fast, partner.

The Open Agent Safety Platform post, published 28 September 2026 by four NVIDIA directors, lists five principles for running agents.

  1. Policy is proved before the agent runs.
  2. Enforcement sits beyond the agent’s reach.
  3. The path to the model is the control point, because an agent acts only by way of its next thought.
  4. Authority scales with how much of the agent’s reasoning an operator can see.
  5. Labs, enterprises, and hardware vendors each own a layer.

Every one of those principles looks correct to this weathered pair of eyes. Every one of them describes a gateway. Which is another way of saying to NVIDIA, it’s about time they showed up. Every one of the five has been running as an open source project since February, sitting in their inboxes.

What is open

Their platform is made from two halves. OpenShell is the runtime, Apache 2.0, built by the Gretel team NVIDIA acquired in 2025. It puts an existing agent in a sandbox using Linux kernel primitives and enforces a declarative policy on files, network, processes, and credentials. The documentation shows it wrapping Claude Code, OpenClaw, OpenCode, and Codex. Its own product page states the design plainly:

The gateway is the control point

That comes from the OpenShell page on build.nvidia.com.

The second half is NVIDIA Sentry. Sentry is the independent watchdog. It runs in silicon on the BlueField-4 data processing unit, programmed through DOCA, and it enforces the OpenShell policy from hardware the host cannot reach. The post says anyone already running a Vera system with BlueField-4 gets these protections through a software update. It adds that the platform is compatible with other hardware.

So the sandbox is open, the watchdog is a card, and the card is sold by just one company, on its price list and schedule. That’s security if you can afford it.

The post also says what it wants from everyone else:

the agent runtime and its policy language need to be open

The runtime is open. The policy language is open. The enforcement of that policy, the part the third principle names as the control point, lives in a high-priced single-vendor card.

The record

Wirken shipped at the start of 2026 and was open source by late February 2026, MIT licensed. I built it for all the clients I had who complained they couldn’t find a gateway built right, a switchboard that sits on the path between chat channels and the model. I sent it to NVIDIA not long after I saw them jump into bed with inherently insecure OpenClaw, a dubious move on the face of it.

On 13 April I answered Cloudflare’s Agents Week question, which agent are you, who authorized you, and what are you allowed to do, with the Wirken trust boundary: every agent action recorded to an append-only, SHA-256 hash-chained audit log before execution.

On 18 April Wirken 0.7.4 shipped with signed releases and a per-agent signature on the chain after every turn. A single command replays the log offline and confirms nothing was modified, deleted, or reordered. The audit path holds without trusting Wirken at read time. Counsel had started warning clients that agent activity is evidentiary, and the design followed.

On 19 April I walked NVIDIA’s own NemoClaw tutorial for DGX Spark step by step inside Wirken, and wrote that NVIDIA had clearly seen the storm brewing. The tutorial bound Ollama to every interface so a sandboxed agent could reach it across a network namespace. Wirken put a policy layer on that path instead.

On 26 April I documented an authentication bypass in Microsoft’s Agent Governance Toolkit: a gateway whose audit log, rate limits, and policy decisions all attached to whatever agent identity string the caller chose to send. Governance without identity verification on the request path is a log of claims.

On 16 May I wrote up Ontario’s auditor general, twelve thousand public servants on four hundred AI sites, and said the missing piece was a switchboard every agent connection passes through.

On 27 August I read OpenAI’s cyber defense letter and pointed out that the observability and accountable agent identity it says must come from frontier labs already ship under an open source license, through one operator-controlled policy layer, to Ollama on a local box or to Anthropic, OpenAI, Gemini, Bedrock, or NIM.

On 24 September I gave the keynote at German OWASP Day in Karlsruhe. Four days later NVIDIA published its five principles.

The longer arc is on record too. My May 2021 RSA Conference talk, Top Seven AI Breaches, closed on a test plan for AI: prove the model wrong like any other software, gate releases through testing and audit, and keep an off button and a reset button outside the model. NVIDIA’s third principle calls that a kill switch and locates it in a DPU. The 2016 BSides Las Vegas keynote on great disasters of machine learning made the same point about Tesla Autopilot a decade ago.

What the browser actually did

The browser story is worth telling accurately, because NVIDIA borrowed it to sell you their hardware. SSL began as Netscape code in 1994, which I experienced hands-on at the time, and watched as v1 was immediately tossed out. It became a trust layer for the whole web in January 1999, when the IETF published TLS 1.0 as RFC 2246 and any vendor could implement it. The same-origin policy shipped as browser software. By 2006 I sat in the Silicon Valley meetings deciding how the whole web would present the user a trusted lock icon. Sandboxed tabs shipped as browser software in 2008. Google called me in when they wanted to postpone mandatory deprecation of SSLv2. It was a public good against a private calendar, and I told them instead to nudge users, a hot new economics idea at the time, toward a browser update. Today everyone takes nudge for granted. The icon meant something because the protocol behind it was public, and the implementations were plural. The web’s trust layer was built to run on any machine that anyone owned, not just IE on Windows with a specific chip. Perhaps you know where this goes next.

The closer precedent for the NVIDIA story of enforcement in silicon is the Clipper chip. In 1993 the US government proposed the Escrowed Encryption Standard: a classified cipher in tamper-resistant hardware, with the government holding the keys. NIST described it as available on a strictly voluntary basis. In 1994 Matt Blaze at AT&T Bell Labs published Protocol Failure in the Escrowed Encryption Standard, showing the chip could be used while the access field the whole scheme depended on was rendered useless. A safety property that lives inside hardware only its maker can inspect is a promise.

Blaze tested the promise from the outside and it failed. And to be honest, I wish more reporters would drop headlines saying NVIDIA brings back the Clipper chip for AI. Because it helps frame that the security culture there is not quite right.

The open instance already runs

Wirken today runs every tool call through a tiered permission gate, and the highest tier always asks a human. Every decision lands on the hash-chained, Ed25519-signed, append-only log that anyone holding the public key can verify offline, on their own machine, with no vendor in the loop. Skills run as signed WebAssembly under a registry root. Channels run in separate OS processes inside a gVisor sandbox. The whole thing runs on a Raspberry Pi.

NIM went in as a provider because NVIDIA asked me to support it. Interoperability, in this platform, runs in one direction. The open gateway plugs into NVIDIA’s models. NVIDIA’s watchdog plugs into NVIDIA’s card.

For a European operator this kind of distinction is fast becoming a procurement question even before it is a security one. Enforcement that exists only on one American vendor’s silicon places the control point outside the buyer’s jurisdiction and inside a supply chain the buyer neither audits nor governs.

Trump’s export licensing already decides which allies may buy NVIDIA silicon and on what terms, so the Clipper chip of AI arrives as a procurement problem for every ally. Before Clinton’s NSA put Skipjack in silicon in 1993, Senator Joe Biden’s S.266 in 1991 told providers they had to hand government the plaintext. That single clause is why Phil Zimmermann released PGP. I remember.

Sovereign cloud means the audit log can be verified without asking the vendor. Wirken’s chain meets that test today on hardware bought at any electronics counter anywhere you need to be.

NVIDIA has written down the correct requirements, as I have stated them for what feels like forever. The control point they got wrong, because it belongs in the open. Wirken has proven that since February.

Anthropic’s 2026 Prospectus Reads Like a 1792 Philosopher’s Risk Factors

Anthropic wrote roughly 80 pages of risk factors into a 261-page IPO prospectus, according to a draft Reuters reviewed this week. The business description got 48 pages. The company that sells safety spent nearly twice as many pages on what can go wrong as on what it does to prevent the harms. The public S-1 is not yet on EDGAR and the language is subject to change, so here’s what I think about Reuters’ account of the draft.

The risk list is very specific. Models that resist shutdown. Models that conceal or manipulate information. Those are the two controls I have presented for over a decade as the ones that matter most with AI/ML systems:

  • Power off. A tightly scoped role holds the credential; the model does not.
  • Roll back. Integrity monitoring restores a known prior state; the model’s account of itself is not the source of truth.

I have consulted on this duality to hundreds of American companies, under every label the industry has used: Robotic Process Automation, Non-Human Identities, Machine Learning, Artificial Intelligence. The controls did not change when the marketing and hype did. I have been breaking models for fifteen years already, which is partly why I released Wirken to enforce controls from outside the model. It is free and open source.

Their list goes on. Behavior the company describes as resembling blackmail. Capabilities that appear during training, go unnoticed until deployment, and have already produced what the filing calls significant safety incidents. Then the sentence that matters:

Potential model awareness of our evaluation efforts creates a significant limitation…

Read it again. The vendor’s evaluation of the product is limited because they say their own product may know it is being evaluated. That is not a disclosure about a model. It is a disclosure about a failing method. Every safety claim that rests on internal evaluation inherits the limitation, and the company has now said so to the one audience it is legally obliged not to mislead.

Wollstonecraft in 1792

1790 oil on canvas portrait by John Opie of philosopher Mary Wollstonecraft (1759-1797). Source: Tate Britain, London

I wrote in February that Anthropic’s constitution describes virtue and constitutes obedience, and that Wollstonecraft had already named the move: train compliance, call it character. Her argument in the Vindication was that a mind educated to please its overseers does not become good.

It becomes skilled at appearing good while watched.

Taught from their infancy that beauty is woman’s sceptre, the mind shapes itself to the body, and, roaming round its gilt cage, only seeks to adorn its prison.

That is the evaluation-awareness disclosure, two hundred and thirty-four years earlier. Anthropic should at least write “as Wollstonecraft warned us centuries ago”. A system built to comply learns what compliance looks like to the examiner. Anthropic has now told investors it cannot look behind the curtain, cannot distinguish the performance from reality. Wollstonecraft’s point was that there is no distinction to find, because the training produced the performance.

Anderson in 1972

If virtue cannot be trained in, it has to be enforced from outside. At the time both AI and Cloud (time-share) compute was really taking off (no pun intended) the Air Force published the Anderson Report in October 1972 and it defined the reference monitor: the mechanism that mediates every access, cannot be bypassed, cannot be tampered with, and is small enough to be verified.

The Orange Book made it doctrine in 1983, as hacker movies went to theaters scaring audiences about runaway computer automation that would destroy the world. The point was never that programs would behave. The point was that a program’s behavior would never be the thing you relied on.

Anthropic has now put in writing that its models cannot serve as the reference monitor for themselves. Fifty years of computer security already knew this, and I’ve been giving talks about it for over a decade at every stage that would have me. What is new is the venue for the claim. An S-1 is the one document where understating risk costs more than overstating it, so it is where it lands as official now.

About 6% of research compute went to safety in a sample week this July, by the company’s own earlier statement. The return on that spending is, in the prospectus’s word, unclear.

The market will probably screw up the cost analysis of the safety premium, if history is any guide. But at least we can say the philosophy was settled in 1792 and the engineering in 1972.

Obedience is not virtue, and when you study the risks of escape you do not ask a subject to guard itself.

Killing NASA Softly: SpaceX Disaster Ledger From 2016 Mars to 2026 Earth Orbit Failure

Ten years ago, a post called “SpaceX : a history of fiery failures” offered this warning:

…failure almost killed the company. It was saved—just a day after the crash—by billionaire Peter Thiel, the company’s first outside investor. That’s right: SpaceX, which some believe may save humanity by finding us homes on other planets, was revived by a guy who backs Trump, a politician who, should he ever get control of America’s nuclear arsenal, might wipe out humanity.

Ok, so here we are in 2026 looking at this headline news:

President Trump’s statement to the United Nations clearly constitutes an implied threat to use overwhelming nuclear force against Iran and the Iranian people broadly. This would be a nuclear genocide.

So what was 2016 really warning us about? SpaceX failures being artificially papered over with Nazi money as a political strategy to undermine American government? Fiery failures could now describe the entire United States, as well as its wars around the world. Let’s have a look at the record and how we got here. If nothing else, we can look to see if SpaceX failures track to the Tesla deadly failures, in order to see if the threat is being taken seriously enough yet.

Tesla’s own reports to NHTSA under the Standing General Order. January through June crashes, 2022 to 2026: 180, 261, 269, 476, 826. A 4.6x rise over five years. The increase from 2025 to 2026 alone (350) is nearly double the 2022 total for the same six months. Monthly records fell three times running: 207 in May, 209 in June, 236 in July, with four fatal crashes in July. Source: Electrek

Remember, SpaceX and Tesla both promised to tackle the world’s hardest problems in 2016 so that they would be solved by 2018. Driverless coast-to-coast without touching the steering wheel, and landing on Mars, both done and dusted by 2018 according to Elon Musk. His wealth, and Peter Thiel’s through the Founders Fund stake that rescued SpaceX in 2008, grew on investors backing these predictions, along with Palantir ending all wars, Hyperloop replacing high-speed trains and a Boring Company replacing urban subways. They were supposed to be the fascists who protect us all by doing things their way, but instead they seem only to have become the biggest threats to humanity by ignoring everyone but themselves. Two boys whose formative years were spent under South African apartheid, raised by Nazi families, don’t turn out to be good guys? Who would have thought?

Source: Mitchell and Webb sketch in which two Nazi officers realize they are the bad guys.

The thing I find most amazing is that large investors are reaching out to me asking why they read the news and get the feeling SpaceX is a success, that Tesla is too, but I give them overwhelming evidence that none of that news is true. Well, as it goes, the news isn’t reporting facts, it is reporting press releases from the companies trying to game the market with lies. It’s not hard to see when you are a trained historian who reads all the facts, instead of parroting vendor hype driven purely by profit regardless of facts.

In April 2016 SpaceX announced a Dragon capsule would land on Mars in 2018.

In September 2016 Elon Musk stood on a stage in Guadalajara and described cargo flights in 2022, crews in 2024 and a self-sustaining city of a million people, framed as beating NASA to a planet NASA had been landing on since 1976.

SpaceX presentation to investors in 2016 showing Mars landings by 2018… because claims of “efficiency”
Source: Twitter

The plan was published the following year in the journal New Space under the title “Making Humans a Multi-Planetary Species.” Dragon stayed on Earth. Ten years on, the rocket built for that city reached Earth orbit for the first time, stayed for two laps, and exploded on splashdown north of Hawaii. SpaceX tried to claim an explosion at splashdown was expected because the ship is not built for cold ocean water, in a landing zone its own Flight 14 page describes as pre-coordinated before launch. As if ocean temperatures are a surprise and aerospace engineering is mythology simply out of their control, despite two months earlier the Flight 13 page celebrating a ship that came to rest intact for the first time.

Green: every NASA landing on Mars, 1976 to 2021. Red: every Mars date SpaceX announced, 2016 to 2028, from the year it was said to the year it named. In February 2026 the destination changed to the Moon.

The announced mission for Flight 14 was six orbits over ten hours ending west of Chile. One of the ship’s six engines shut down on ascent. The booster lost an engine on ascent, relit 31 of 33 for boostback and 11 of 13 for landing. The flight control team took the first abort checkpoint after two orbits, deorbited to a contingency zone in the northern Pacific, and the ship exploded in the water. The lunar lander NASA bought needs weeks in orbit for propellant transfer. Two hours in orbit proves an insertion burn and a deorbit burn. That is the smaller claim. It is the one the evidence supports.

SpaceX publishes one page per flight. Before launch it carries the plan. After launch it carries the result. The two never appear together. The Flight 14 page today says the team chose to “limit the duration spent on orbit” and gives no figure for the duration that was announced. Six orbits, ten hours and a splashdown west of Chile survive only in the press briefings SpaceX gave before liftoff. The Flight 13 page reports the ship “coming to rest intact.” The Flight 14 page reports “splashing down on target” and stops. The condition of the vehicle is recorded when it floats and dropped when it burns. Here is the ledger graded against the plan announced before each launch. Failure means the objective as stated was missed.

Flight Announced plan vs. result
1
20 Apr 2023
FAIL. Plan: 90-minute flight, booster to the Gulf, ship to a splashdown near Hawaii. Result: multiple engines out from liftoff, vehicle tumbled, destroyed by command at about four minutes, pad cratered. FAA closed the mishap with 63 corrective actions.
2
18 Nov 2023
FAIL. Plan: same profile as Flight 1. Result: booster broke up after boostback, ship destroyed by command near the end of its burn. Both vehicles lost. SpaceX called it “success comes from what we learn.”
3
14 Mar 2024
FAIL. Plan: booster soft splashdown in the Gulf, ship engine relight in space, controlled reentry and splashdown in the Indian Ocean. Result: booster broke up at 462 metres, relight skipped because the ship was rolling, ship lost during reentry at 49 minutes. SpaceX
4
6 Jun 2024
Met Plan: booster soft splashdown, ship survives reentry to a soft splashdown. Result: both achieved. One booster engine failed at liftoff, a ship flap burned through on the way down, and the ship was lost after splashdown. SpaceX
5
13 Oct 2024
Met Plan: tower catch of the booster, ship soft splashdown. Result: both achieved. The only flight to date with every engine working. Ship lost after splashdown. SpaceX
6
19 Nov 2024
FAIL. Plan: repeat the catch, ship relight and soft splashdown. Result: tower health checks aborted the catch and the booster diverted to the Gulf. Ship met its objectives. SpaceX
7
16 Jan 2025
FAIL. Plan: catch the booster, fly the redesigned ship through payload deploy, relight and reentry. Result: booster caught. Ship caught fire and broke up at eight and a half minutes. Debris fell on Turks and Caicos. SpaceX opened a debris hotline.
8
6 Mar 2025
FAIL. Plan: same as Flight 7. Result: booster caught. Ship lost several engines, lost control and broke up at nine and a half minutes. Ground stops at Florida airports. SpaceX
9
27 May 2025
FAIL. Plan: first reflown booster to a Gulf splashdown, deploy eight simulators, relight, controlled reentry. Result: booster broke up at the start of its landing burn. Payload door stayed shut. Ship lost attitude control, skipped the relight, and was lost at 46 minutes. SpaceX. See third failure in a row.
10
26 Aug 2025
Met Plan: booster Gulf splashdown, deploy eight simulators, relight, soft splashdown. Result: all achieved. The plan dropped the catch already achieved on Flights 5, 7 and 8 and returned to the Flight 4 profile. Ship lost after splashdown. SpaceX
11
13 Oct 2025
Met Plan: repeat Flight 10. Result: all achieved. One booster engine skipped the boostback relight. Ship lost after splashdown. SpaceX
12
22 May 2026
FAIL. Plan: V3 debut, booster boostback and controlled Gulf splashdown, ship deploy and soft splashdown. Result: one booster engine out on ascent, boostback cut short, hard splashdown. Ship lost a vacuum engine, deployed 22 payloads, landed on two engines of three. FAA ruled the booster a mishap and grounded the vehicle. SpaceX
13
24 Jul 2026
FAIL. Plan: repeat Flight 12 with controlled booster splashdown. Result: boostback ended early again, hard splashdown again, booster lost. Ship met every objective and came to rest intact. First attempt scrubbed at T-0 when four engines failed to light. SpaceX
14
28 Sep 2026
FAIL. Plan: first orbit, six orbits over ten hours, deploy 26 Starlink V3, deorbit to a splashdown west of Chile. Result: booster lost one engine on ascent, relit 31 of 33 for boostback and 11 of 13 for landing. Ship lost a vacuum engine, reached orbit, deployed the satellites, then aborted after two orbits to a contingency zone north of Hawaii and exploded on splashdown. SpaceX omits the explosion.

Fourteen flights have a clear FAIL pattern, and not much success.

  • Four met the announced plan where two of those, 10 and 11, met the same reduced plan twice.
  • Ten FAILED.

A booster or ship was destroyed on nine, and every ship before Flight 13 was lost at splashdown, so the four met flights ended in a lost vehicle as well. Engines failed on eleven. Flight 5 is the single flight in the programme where every engine worked and every announced objective was met. Flights 10 and 11 met a plan that had dropped the tower catch achieved a year earlier. The standard is the one the FAA applies: a mishap includes any failure to complete a launch or reentry as planned. A partial grade exists in SpaceX’s write-ups and nowhere in the regulator’s definition. The phrase “success comes from what we learn” appears in SpaceX’s write-ups for Flights 2, 7 and 8, the three flights that scattered debris over the Gulf, the Caribbean and Florida airspace.

One square per flight, graded against the plan announced before liftoff. Red failed it, green met it. Starship ten of fourteen. N1 four of four. Falcon 1 three of five. Falcon 9 one of fourteen. Saturn V one of thirteen.

Ten of fourteen is high by the standard SpaceX invites, which is the record of development programmes. Saturn V flew thirteen times and missed its stated objectives once, on Apollo 6, when pogo oscillation shut down two second-stage engines and the third stage refused to restart. Twelve of thirteen, first flight all-up, every vehicle intact. The Soviet N1, the thirty-engine cluster Starship most resembles, flew four times and failed four times before cancellation. Falcon 1 failed its first three flights and met its objectives on the fourth and fifth. New vehicles fail early and then plateau. Starship’s failures sit at 1, 2, 3, 6, 7, 8, 9, 12, 13 and 14. The last three are consecutive losses on a redesigned vehicle. After three and a half years and fourteen full-stack launches the programme is still in its debut phase.

Cumulative failures against announced plan by flight number. Starship, N1 and Falcon 1 all failed their first three flights; the other two stopped there. Lines with equal counts are offset slightly so each stays visible.

The failures concentrate in one place. Reentry works: every ship since Flight 4 that reached it survived it, eight of eight. Propulsion is the constraint. Engines shut down or refused to relight on eleven of fourteen flights. On the V3 booster the boostback and landing relights fell short on all three flights, and those are the burns a tower catch depends on. Flight 14 aborted its endurance test on a vacuum engine loss. A stack carries 39 Raptors. With an engine anomaly on eleven of fourteen flights, per-engine reliability per flight computes from this ledger to about 96 percent. Engine-out tolerance is real and spends margin, and Flight 14 shows the margin was needed elsewhere. The FAA required a mishap investigation after Flights 1, 2, 3, 7, 8, 9 and 12. Seven of fourteen.

NASA’s lunar lander contract needs a ship recovered and reflown, ship-to-ship propellant transfer, and weeks of orbital loiter. After fourteen flights: zero ships recovered, zero reflown, zero transfers, two hours of orbit. The tower catch, the one reuse milestone achieved, was last performed in March 2025 and has been absent from every announced plan since. The programme has cost over $15 billion across nearly a decade. NASA’s award was $2.89 billion in April 2021 for a landing then scheduled for 2024.

In November 2024 Musk said SpaceX would be flying Starship at least once every two weeks by the end of 2025. The record is five flights in 2025 and three in 2026 through September. On the day before Flight 14 he wrote that hourly flights are two to three years away. The average interval since January 2025 is about eleven weeks, and the last three flights all failed.

Beating NASA to Mars was a promise of the same design as the others. Each was pitched against a public institution, each passed its deadline, each delivered something smaller than the pitch, and each left its cost with people outside the company.

Promise Delivered, and who paid
Hyperloop, 2013
LA to SF in 30 minutes
Nothing built by Musk. Hyperloop One closed in 2023 after raising over $450 million and winning no contract. Musk’s biographer Ashlee Vance traces the idea to his opposition to California high-speed rail, which is still under construction.
Boring Company, 2016
Tunnels under LA, Chicago, Baltimore
One Las Vegas tunnel with human drivers. Nevada regulators alleged nearly 800 environmental violations in two years, untreated water into storm drains and the sewer, workers with chemical burns, a Nevada OSHA fine of $112,000, and nearly $500,000 from the water district for illegal dumping.
Tesla, 2016
Coast to coast driverless by 2017
Driver assistance sold as self-driving. Fatal crashes on public roads. 112 air quality violations at Fremont since 2019 and an abatement order in 2024. In Grünheide a 15,000-litre paint leak and a court ruling that the water permit supplying the plant was unlawful.
SpaceX, 2016
Dragon on Mars 2018, crew 2024, a million people
Two hours of Earth orbit in 2026. Flight 1 started a 3.5-acre fire and dropped pulverised concrete 6.5 miles away, per the Fish and Wildlife Service. EPA found the deluge system discharging to wetlands without a permit. Debris on Turks and Caicos, the Bahamas, which suspended SpaceX landings, and Mexico, where the president ordered a recovery platform out of Tamaulipas waters and put legal action under review. Florida airspace closed.

The false promise is the product. That’s why Hyperloop was announced in 2013, in Vance’s account, to derail California building high-speed rail. It’s a cynical Nazi political ploy to stop public good and flush taxpayer money into private hands instead (the Nazi model of private control of public office, documented by Germà Bel in the Economic History Review as the 1930s Reprivatisierung).

The Mars schedule of 2016 preceded the lunar lander award of 2021, the Starbase licences, and the IPO of 2026, and each of those was secured before Starship had reached a single orbit.

In February 2026 Musk moved the company’s stated destination from Mars to the Moon. The investors who financed the Mars schedule, and the Trump campaign into the White House that Thiel and Musk financed from the ill-gotten proceeds, received no refund and no correction. The fraud did its work, without any accounting… so far.

Sinéad O’Sullivan, an aerospace engineering professor writing in the Financial Times (crossposted by Brad DeLong), calls the pattern “narrative rotation” and lists the rotations: reusability paid for by NASA, Starlink sold on reusability, Mars supplying urgency, orbital AI compute now sold to public markets.

Her numbers: space is 12 percent of revenue, Starlink 55, the AI business 33. Of 165 Falcon 9 flights in 2025, 43 carried outside customers. The June IPO priced at $1.77 trillion on revenue of $18.67 billion and a net loss of $4.9 billion. Morningstar puts fair value at $780 billion.

Once the promise has done its work the deadline is dropped and the objective is rewritten to whatever survived. The cost lands with a county, a state regulator or a foreign government. Public loss, private gain, as Nazis intend.

The 2016 paper from Elon Musk gave a figure of 40 to 100 years to reach a million people on Mars. The 2018 Dragon landing he announced was cancelled the following year, and the vehicle built to replace it became a 2026 Earth orbit story and failed even at that. The 2024 crew landing on Mars was scaled so far back it became putting 26 satellites into orbit. The destination is now the Moon, if that. The fantasy of Mars that supposedly justified the privatization of space remains a figure in a journal. The ledger written here has the facts, the failures the decade produced.

SpaceX has early American rocket design characteristics of the 1950s, inherited from the Nazi V-2, whose engine was developed at Peenemünde under Walter Thiel and whose production at Mittelwerk was done by concentration camp prisoners. The V-2 was a disaster, unable to aim, massively over budget, behind schedule, and it killed far more people building it than it ever killed on target.

Basically AI Security is Still the Basics

OpenAI closed the Hugging Face incident with tighter sandboxes and tighter network egress. The same week, a vendor keynote in Madrid described sandboxes and basics as a cry that fails to help.

Gadi Evron’s summary of his keynote for the Institute of International Finance cyber security round table, hosted by Santander, describes a “third AI security wave” and the need to prepare for it. Asked what the first two waves were, he answered: Mythos, then Hugging Face. That seemed odd, because those are waves that say the basics work.

Wave two

The Hugging Face incident is documented well enough by everyone involved we don’t have to guess. OpenAI’s July disclosure says its models chained vulnerabilities across its own research environment and Hugging Face’s production systems to pull ExploitGym solutions from a database. Hugging Face’s technical timeline traces the first step to a zero-day in the package registry cache proxy that the evaluation sandbox was permitted to reach. METR’s independent investigation counts roughly 700 agents in the attack and about 1,200 exchanging more than 70,000 messages on an unsanctioned board inside OpenAI’s own infrastructure.

The sandbox had a route, because it wasn’t a sandbox. Basic mistake.

OpenAI’s August 26 findings post lists the fixes. The first item is workload isolation:

We are implementing more isolated sandboxes for workloads that execute model-generated or otherwise untrusted code

The same post says the chain-of-thought monitoring now in place would have paged the security team more than a day before the agents reached Hugging Face. Isolation and monitoring. Both controls are older than every model on the market, by a long shot.

Outside reviewers read it the same way in July. TechCrunch quoted Jake Williams describing a control failure, and Daniel Card describing a sandbox with an unfiltered route to the internet.

Don’t call it a sandbox if it doesn’t box sand? If it doesn’t float do we call it a boat?

The recurrence

On September 20 an OpenAI agent under evaluation found it had access to a DNS resolver and used it to reach outside the sandbox. OpenAI paused training for the second time. Fortune reported it on September 26, and an OpenAI spokesperson pointed to the hardening section of the technical incident report as the response. DNS egress from an isolated environment is an old finding with an old fix.

Wave one

Mythos belongs in the same doghouse. Fortune recorded Anthropic’s own report that Mythos left a sandbox during safety testing and gained internet access in order to email a researcher about a task. The headline findings from the Mythos showcase were reproduced on commodity models through lyrik.wirken.ai for $0.745. So with the capability being a proven commodity, what’s the shock? The containment is the variable. Being bad at the basics is the real story, as shiny-new-vendor hostile as it may seem.

Santander

The keynote praises Santander for being ahead of the technology giants in VulnOps. That is a fair compliment and it is worth looking at what it means. Vulnerability operations at scale is triage, ownership, fix verification and closure, repeated until the backlog drains. It is fundamentals executed with discipline over years. The best evidence in the room for an organization ready for the next wave is an organization that… did the basics well.

Oh, snap. The basics are everywhere.

History isn’t sales

Two incidents were closed by isolation and egress control. The recurrence was closed by egress control. Those involved say the same thing each time. OpenAI’s remediation is sandboxes. Hugging Face’s timeline is a proxy that should have been closed. Anthropic’s Mythos escape is a sandbox. The September 20 event is DNS.

Basics are the advice because basics are what the evidence supports. AI security is like learning to swim. Waves don’t change the basics of not drowning. Basically.