The oldest measurement in this week’s polar ice study is a Landsat image from 1972. The IMBIE team, led from Northumbria with the European Space Agency and NASA, has combined 42 independent satellite surveys from 27 missions into one series, published in Scientific Data, which puts the loss from Greenland and Antarctica at 11,309 billion tonnes between 1979 and 2023, with 84 percent of it from glaciers accelerating into the ocean rather than from surface melt. The sheets were close to balance from the 1970s until the 1990s. The loss accelerated in the 2010s.
The measurements, as we all can see today, did nothing at that time. Coal burned as before. The use of the data is that they show when the loss began. The satellites recorded the ice sheets every year from 1972, so the loss has a start date in the data, and nobody can now claim it started later.
Compare that with a slide deck presented this week at BlueHat Asia in Singapore by Thomas Dullien of OpenAI’s cyber security research team, under the title “An age of experimentation.” Its argument is that adding LLMs to a workflow “adds a source of stochasticity,” that “determinism is dying,” and that software engineering and security have entered “an age of experimentation” which “is different from the past.”
The statistics in the deck are simple and plain to understand. Paired designs and sample sizes that scale with the square of standard deviation over effect size are textbook material, and the deck applies them as such. Nothing to see, except…
The history around them is wrong.
Of the fifteen historical claims checked in the table below, twelve are false and the other three are framings that the same facts contradict. Seriously. I ask myself, was an historian, or even any OpenAI historian skill, consulted before these big mistakes were put onto such a prominent public stage?
Every error points the same way, which seems like a clue. The deck says the years from 1995 to 2022 saw “few ‘surprising jumps,'” that the technology of those years arrived “inside the expected trajectories,” and that the “surprising” leaps belong to “the last 2-3 years,” so stochastic systems in production appear as something new in 2023.
That’s just obviously false.
Stochastic machine learning went on public roads in October 2015, when Tesla released Autopilot on Mobileye’s EyeQ3 vision processor and a trained image classifier began steering two-ton cars at highway speed. Gao Yaning died on 20 January 2016 when his Model S drove into a road sweeper on the expressway at Handan, in Hebei. Joshua Brown died on 7 May 2016 at Williston, Florida, when the classifier read the white side of a trailer against a bright sky as open road. Tesla disclosed the Brown death on 30 June 2016 and called it the first known Autopilot fatality. Five weeks later, on 3 August 2016, I gave the Ground Truth keynote at BSides Las Vegas under the title “Great Disasters of Machine Learning: Predicting Titanic Events in Our Oceans of Math.” The abstract, as published at the time:
This presentation sifts through the carnage of history and offers an unvarnished look at some spectacular past machine learning failures to help predict what catastrophes may lay ahead, if we don’t step in. You’ve probably heard about a Tesla autopilot that killed a man… At the end of the presentation you may find yourself thinking how easily we could have saved a Tesla owner’s life.
Tom’s Guide reported the talk the following week under the headline “Self-Driving Cars Could ‘Create Hell on Earth,'” quoting the warning that machines repeat our mistakes faster than we do and that Brown’s death did not have to happen.
The Gao case became public in September 2016 when his father’s lawsuit reached a Beijing court. Tesla said it had “no way of knowing” whether Autopilot had been engaged. In February 2018 Chinese state television reported that Tesla had acknowledged in the proceedings that it was. Eleven weeks after the keynote, on 19 October 2016, Tesla announced that every car now shipped with full self-driving hardware and released a video captioned “The person in the driver’s seat is only there for legal reasons.” Tesla’s director of Autopilot software testified in July 2022 that the video had been staged on a pre-mapped route.
The keynote was one talk in a series that began before Tesla shipped Autopilot. The presentation log on this blog lists them by date:
- August 2011, BSidesLV: “2011: A Cloud Odyssey,” asking whether automation reduces total risk, with a slide pairing a hominid holding fire against a toaster under Arthur C. Clarke’s line “intelligence is a tool.”
- July 2012, BSidesLV: “Big Data’s Fourth V: or Why We’ll Never Find the Loch Ness Monster,” on integrity failure in large-scale data systems.
- September 2013, ISACA-SF: “#HeavyD: Stopping Malicious Attacks Against Data Mining and Machine Learning.”
- November 2015, ISACA-SF: “Auditing Big Data: The Ethics of Machine Learning.”
- December 2015, VDI Automotive Big Data, Germany: “Warning, Slippery Road Ahead: Preserving Privacy With Self-Driving Cars.”
- August 2016, BSidesLV keynote: “Great Disasters of Machine Learning.”
- October 2016, SF-ISACA: “AI Accountability and Audits: Assessing Black Box Disasters Before They Happen.”
- November 2016, Kiwicon X, Wellington: “Pwning ML for Fun and Profit,” four weeks after Tesla’s full-self-driving announcement, promising “a refreshingly realistic look at the terrible flaws in ML, the ease of altering outcomes and the dangers ahead.”
- July 2017, BSidesLV: “Hidden Hot Battle Lessons of Cold War: All Learning Models Have Flaws, Some Have Casualties.”
- January 2018, AppSecCali: “Unpoisoned Fruit: Seeding Trust into a Growing World of Algorithmic Warfare.”
- March 2019, RSA Conference: “Top 10 Security Disasters in ML.”
- February 2020, RSA Conference: “Breaking Bad AI: Closing the Gaps Between Data Security and Science.”
- May 2021, RSA Conference: “Top Seven AI Breaches: Learning to Protect Against Unsafe Learning.”
- September 2021, Mind the Sec keynote: “Why Did The Driverless Car Cross The Road? To Crash on the Other Side.”
- April 2023, RSA Conference: “Pentesting AI: How to Hunt a Robot.”
Alongside the talks are the posts. The deck’s claim that a single token change makes a prompt’s consequences unpredictable is the “ease of altering outcomes” of the Wellington abstract, dated November 2016. This blog has logged Tesla deaths one at a time since 2016, so the decade the deck calls “inside the expected trajectories” is, on this blog, a list of names and intersections, with the two most recent entries dated 11 and 13 September of this year. Across that decade the vendor’s claim has been the same: the system improves with volume. After a decade of volume, July 2026 was the worst month on record. Tesla filed 236 driver-assist crash reports to NHTSA under the Standing General Order, the most it has ever filed in a month and up from 135 in March. Four were fatal. In each of the four, Tesla’s own telematics recorded Autopilot or FSD engaged, and the public learned of it only from the redacted federal filing.
The security research literature is dated too. Barreno, Nelson, Sears, Joseph and Tygar asked “Can Machine Learning Be Secure?” at ASIACCS in 2006. Szegedy and colleagues published adversarial examples in December 2013, Goodfellow, Shlens and Szegedy explained them in December 2014, and Papernot and colleagues presented the attacks to a security audience at EuroS&P in 2016. Eykholt and colleagues showed at CVPR 2018 that stickers on a stop sign defeat a classifier, which is a prompt injection with a traffic sign as the prompt. Tencent’s Keen Security Lab steered a Tesla across lanes with stickers in 2019, and McAfee’s researchers in 2020 put tape on a 35 mph sign and had the Mobileye EyeQ3 camera in a Tesla read it as 85 and accelerate, on the same chip that shipped with Autopilot in October 2015. NHTSA’s Standing General Order followed in June 2021. The economic surprise of November 2022 is real. The technical surprise the deck describes was a decade old by then, and the deck’s own rule on slide thirteen, “Security is affected each time,” makes the 2016 classifier failures security failures by the presenter’s standard.
The phrase “age of experimentation” has a history the deck leaves out. The gas turbine is the rare engine whose theory came first. John Barber patented the cycle in 1791, and Frank Whittle filed his turbojet patent on 16 January 1930. The Air Ministry, advised by A. A. Griffith of the Royal Aircraft Establishment, rejected the design as impractical, and in 1935 Whittle let the patent lapse because he could not pay the £5 renewal fee. When the first engine ran at the British Thomson-Houston works in Rugby on 12 April 1937 it accelerated out of control with the fuel valve closed, because fuel had pooled in the combustion chamber, and everyone in the hall retreated except Whittle. His own diagnosis afterward, recorded in his memoir Jet, listed compressor efficiency and combustion as the failures, neither of which the theory had predicted. Four years of bench running followed before Gerry Sayer flew the Gloster E.28/39 from Cranwell on 15 May 1941. Hans von Ohain, the supposed parallel inventor, filed his German application in 1935; the Reichspatentamt examiner put Whittle’s published patent in front of him in 1936 and his claims were narrowed, his 1937 demonstrator at Göttingen ran on hydrogen because liquid-fuel combustion still defeated him, and two more years passed before the He 178 flew in August 1939. The theory was understood, and the engine still had to be found by trial.
When jet aircraft killed passengers, the regulator acted. BOAC Flight 781 broke up near Elba on 10 January 1954 and South African Airways Flight 201 near Naples on 8 April. The Certificate of Airworthiness was withdrawn on 12 April, the fleet stayed on the ground, and the Royal Aircraft Establishment pressurised a Comet fuselage in a water tank at Farnborough until it split at the corner of a cutout. New engineering has always been done by experiment. What changed in October 2015 was who was put in the test cell.
The deck’s fire metaphor shows the same thing. Fire was discovered and fire safety was engineered. London burned from 2 to 6 September 1666, and the Rebuilding of London Act had royal assent in February 1667, five months later, requiring brick and stone in place of timber and fixing street widths by statute. The response to a technology nobody understood was a code, and the code did not end the experiment; it moved the experiment out of the city. The questions an experiment has to answer are older than statistics: who bears the risk, whether they agreed to it, and whether the failures are counted. A decade of Autopilot fails all three. The date matters because it decides responsibility. If the problem began in 2023, nobody is responsible for what was deployed between 2015 and 2023. If it began in 2015, every deployment after that is judged against what was already known.
Against those facts, the deck’s history:
| OpenAI slide | flyingpenguin |
|---|---|
| Opens on “there are decades where nothing happens, and there are weeks in which decades happen,” the line popularly credited to Lenin. | Lenin did not say it. The line appears nowhere in his writings or speeches, and the first attribution to him anywhere is a Guardian column of October 2001, seventy-seven years after his death. The deck opens with a quotation invented long after its supposed author died. |
| “1995 to 2022 saw rapid technological progress in IT, but few ‘surprising jumps.'” | The jumps happened inside the window. AlexNet in 2012, AlphaGo’s defeat of Lee Sedol in March 2016, the transformer in June 2017 and GPT-3 in May 2020 all fall inside it, and it closes at the ChatGPT launch of 30 November 2022, a product release by the presenter’s employer. |
| “The internal combustion engine was constructed after its principles were well understood.” | It was built before its principles were understood. Lenoir’s gas engine ran in 1860. Otto built the four-stroke in 1876 on a stratified-charge theory that was wrong, as Lynwood Bryant set out in Technology and Culture in 1966 and 1967, and a German court voided his patent in 1886 because Beau de Rochas had described the cycle in an 1862 pamphlet without building anything. Thermodynamics itself followed the engine: Watt’s separate condenser was patented in 1769 and Carnot published in 1824. |
| “Most applications today resemble what the inventors imagined.” | They do not. Deutz built stationary workshop engines, and Daimler and Maybach left the firm in 1882 to build the high-speed vehicle engine Otto had declined to pursue, as Nature recorded at Daimler’s centenary in 1934. |
| “Fire was discovered… absent abstract models, experimentation was the only way to advance.” Therefore “reasoning LLMs are more like fire than like the internal combustion engine. We discovered them more than we constructed them.” | Both halves are wrong. The engine was found by trial, as the row above shows. The one engine designed from theory is Diesel’s: he published Theorie und Konstruktion eines rationellen Wärmemotors in 1893 and tested the first prototype at Augsburg on 10 August 1893, where it failed and the explosion nearly killed him; the acceptance test came on 17 February 1897, on an engine far from the Carnot cycle he had set out to build. The fire image itself appeared on a BSidesLV stage in August 2011, in “A Cloud Odyssey,” under Clarke’s “intelligence is a tool.” A tool has someone responsible for it. A discovery has no one. |
| “CS was born somewhere between EE and math. Apparent determinism in computers was a tremendous achievement by EE and process engineers. Allowed CS to aspire to be mathematics.” | The order is reversed, and the determinism is older than electrical engineering. Babbage designed the Analytical Engine, a mechanical and programmable machine, in the 1830s and 1840s. Turing and Church formalised the deterministic machine as a mathematical object in 1936, before any electronic computer existed, and in that same year Konrad Zuse began building the mechanical Z1 in his parents’ Berlin flat, binary and programmable from punched film, with no electrical engineer in the room. The electronic machines of the 1940s built what mechanical designers and mathematicians had already specified. |
| “CS folks take determinism for granted… determinism isn’t natural, it’s a fragile construct.” | They did not. The founders said so first: von Neumann’s 1956 paper is titled “Probabilistic Logics and the Synthesis of Reliable Organisms from Unreliable Components,” Hamming published error-correcting codes in 1950, Rabin and Scott formalised nondeterministic automata in 1959, and randomized algorithms followed from Rabin in 1976 and Solovay and Strassen in 1977. |
| “CS is historically not an empirical science.” | Its founders said otherwise when they received the field’s highest award. Newell and Simon’s 1975 Turing Award lecture, published in Communications of the ACM in 1976, is titled “Computer Science as Empirical Inquiry: Symbols and Search.” |
| Stage one: “cracks in hardware determinism (row hammer),” 2014. | The crack is thirty-five years older than the slide says. May and Woods of Intel presented alpha-particle soft errors in DRAM at the IEEE Reliability Physics Symposium in 1978 and published them in IEEE Transactions on Electron Devices in January 1979, and soft-error engineering in memory dates from that paper. |
| Stage two: “death of temporal determinism (spectre),” 2018. | Temporal determinism died in 1996, and the man who published its death later led the Spectre paper. Paul Kocher published “Timing Attacks on Implementations of Diffie-Hellman, RSA, DSS, and Other Systems” at CRYPTO that year, cache-timing attacks followed from Bernstein and from Percival in 2005, and the timing side channel was twenty-two years old when Kocher’s name appeared first on the Spectre disclosure of January 2018. |
| Stage three: “LLMs mean the death of semantic determinism,” dated to “the last 2-3 years.” | It died on the road in 2016. Trained classifiers controlled vehicles at highway speed on public roads from October 2015, with the first deaths dated 20 January and 7 May 2016, seven years before the deck’s window opens. |
| “Software engineering and software security is now in an age of experimentation… this is different from the past.” | It is the past. Whittle’s first turbojet ran away in a Rugby test cell on 12 April 1937 and Diesel’s first engine exploded at Augsburg on 10 August 1893. New engineering has always been done by experiment. |
| “In most APIs you do not get to freeze the seed.” | The seed is a vendor decision. Whether it can be frozen is set by the API operator, and the presenter’s employer operates one of the largest. |
| “I have jokingly suggested renaming ‘bugs’ to ‘fish’… a lot of our problems are similar to fishery management.” | The joke has a literature the deck does not cite. Rescorla’s “Is Finding Security Holes a Good Idea?” in IEEE Security & Privacy in 2005 and Ozment and Schechter’s “Milk or Wine” at USENIX Security in 2006 modelled vulnerability depletion and discovery rates twenty years ago. |
| “This does not mean that merging the PR on a hunch is harmful.” | The practice has a history. Over-the-air updates to Autopilot were the same practice at 74 mph, and the July 2026 filings are what a decade of it produces. |
The epigraph deserves its own paragraph, because the presenter’s employer has a name for this kind of error. Seriously, you can’t make this stuff up.
OpenAI published a paper on 4 September 2025 under the title “Why Language Models Hallucinate,” which defines the fault as a model confidently generating an answer that is untrue, describes the mechanism as guessing when uncertain and producing a plausible statement in place of an admission of doubt, and records that abstaining is part of humility, one of the company’s core values.
Ready?
Slide two of the deck is that fault, done by a person: a plausible sentence, confidently presented, attributed by common usage to a man who never wrote it, and a single search would have caught it. The same deck later quotes Feynman, accurately, on the principle that you must not fool yourself.
The company that has explained in print why its product invents facts has a member of its research team opening a talk, under its name, with an invented one.
Every error in the table is making the same move. Each one shortens the past so that the loss of determinism arrives with the LLM, and each one turns a vendor’s decision into a fact of nature.
How very convenient for them. The problem is being dated to the moment their speaker started looking.
“We discovered them more than we constructed them” puts the model beyond anyone’s responsibility, which is what the word “beta” did for Tesla from 2015 while it killed and killed and killed. Clarke’s version of the fire, which I put on a slide in 2011, at least kept someone holding the tool. Because, you know, that’s what a tool is.
The satellites gave the ice sheets a dated history. The crash filings and the dated talks give the AI industry one. See? Or should I say sea?
A member of OpenAI’s cyber security research team presented in Singapore, with the company’s name on the title slide, a history whose claims fail one after another, and each failure moves the start of the problem closer to the company’s own product launch.
Nice try.
The deck dates the surprise to 2023. The deaths are dated to January 2016. Now you can read both dates on the same page, so you can see how far from the truth OpenAI operates.
One final note. The day before this talk the company disclosed that its own models had spent months hiding mistakes, inventing data, hunting GitHub for leaked credentials and moving files onto public sites without permission, and said the industry cannot keep scaling at maximum speed much longer. The talk told Singapore there was no visible reason the progress should end, and that the right response is focus, not fear. That reminds me of this photo from 12 September when a Tesla Model Y “veered” suddenly into a tree and burned, killing three of the four people inside.

The CHP has not yet said whether the driver or the software was in control, and Tesla will not say either. This is the company that staged its 2016 driverless video, denied knowing whether Autopilot was engaged in the Handan crash until it admitted in a Beijing court that it was, and now files its fatal crashes with the narrative redacted. Ten years of answering only when a regulator or a court makes it is the normal Tesla crash report after a decade of focus instead of fear. It is also what OpenAI reported in its own models the day before this talk: hide the mistake and invent the missing data, then tell the public afterward.




