ThanksNOTgiving: OpenAI Astra is About as Novel as a Turkey

OpenAI’s summer of claims is turning into a fall with a citation problem.

The company called their Hugging Face breach unprecedented, called their Black Hat account of it a watershed, wrote that ten mathematics problems had seen no progress for a decade, let a “secret technique” story about Astra run for a day, and launched Astra claiming the first model to cross its own cyber threshold. And yet every single one of these is not new, because there is an easily proven prior record. In two cases the prior record is proven by OpenAI’s own documentation of it.

Precedented

OpenAI models in July under evaluation had their safety filters removed and were put in a sandbox with internet access disabled. The models exploited a zero-day in the package registry proxy, moved laterally until they reached a node with internet, and then broke into Hugging Face production to cheat the answers to the test. OpenAI’s description of this was “unprecedented.” I wrote the five whys the next day, explaining what really happened. The chain began with a file that was trusted without a reason to be trusted. Cliff Stoll published the genre in 1989. I gave the BSidesLV talk 23 years later, in 2012, on data as the attack vector for AI. The failure had at least fourteen years of public knowledge before OpenAI’s engineers said it was new in their own evaluation harness.

Watershed

OpenAI in August pumped their story at Black Hat. The presentation treated a sandbox design failure as fixed by rebuilding the compromised service in the same wrong way, and so the agents rebuilt their coordination channel within days. Incident response frameworks require root cause removal. OpenAI presented their inability to perform basic security operations, and basic incident response, as a wakeup in the field. The follow-up was no better.

A decade behind

OpenAI published on 1 August ten mathematics results attributed to an internal Astra model. The release claimed these are problems that “have been open and seen no progress on the main result for at least a decade.”

It’s not a claim that is hard to check. Stephen Miller of Yeshiva University found the sphere-packing proof resting on an argument from his own 2016 paper. Francesco Fournier-Facio of Cambridge found the non-sofic group construction assembled from Gábor Kun’s 2016 paper and the 2019 Kun-Thom paper.

OpenAI’s paper actually cites all three. Cohn-Miller 2016 appears once, for a preliminary reduction step, and Miller says the argument the proof hinges on came from that same paper and was presented as the model’s own. Kun and Kun-Thom are cited in the summary of the non-sofic result. It’s just that their PR describing the paper erased the prior work entirely. Scientific American ran a scathing indictment of OpenAI five days later. Miller called the pattern systematic and put it under research misconduct.

OpenAI edited their page to say instead each result “resolves or makes substantial progress on a long-standing open problem,” under the original date, with no note about being caught.

The same release carries a section on responsibility to the mathematical community. It argues that attribution should reflect how a result was produced, and that claiming human authorship for an AI-generated proof would misrepresent the work.

Sheesh.

Credit between the system and OpenAI’s own staff is settled in that paragraph. Credit to the 2016 and 2019 authors was handled by the sentence above it. An OpenAI spokesperson told Scientific American the company meets the standards expected of human mathematicians. A human mathematician who submits a proof on a 2016 argument credited only for a preliminary step, and describes the field as stalled for ten years, is facing misconduct charges.

Open secret

Four weeks later The Information reported a “secret technique” inside Astra. I traced it the same day: recurrent depth, Graves 2016, Dehghani 2018, Giannou 2023, Geiping 2025 with released weights, and Nanbeige shipping it under Apache-2.0 in July with forty thousand downloads a month. OpenAI, Anthropic and Google DeepMind had cited the Geiping paper by name in their July 2025 chain-of-thought monitorability statement as a risk to document. In that case OpenAI’s own chief scientist publicly rejected the novelty claim being made.

Opposite of first

On 3 September OpenAI launched GPT-6 Astra as the first model to meet the critical cybersecurity threshold of its own preparedness framework. Easy to say, of course, when the threshold is OpenAI’s to say about OpenAI. But everyone knows Anthropic gated Mythos for cyber capability in April. The launch post reports two zero-day vulnerabilities found during evaluation. The July incident report had already described the models chaining zero-days to leave the sandbox.

So vain, so lame

Five claims in ten weeks. Five flops.

One came from reporters working from an anonymous source, and the vendor’s own scientist loudly disowned it. Four came from OpenAI.

The July incident was a repeat of fourteen years of published defense literature. The August paper was 2016 and 2019 mathematics that the paper itself cites. The September architecture was 2016 machine learning that OpenAI itself cited back in 2025.

The company’s PR calls everything new, even when the company’s own documents record earlier work.

OpenAI states the mathematics manuscripts were prepared by humans working with the model. Miller’s 2016 preprint with Henry Cohn, arXiv 1603.04759, is dated March 2016 and sits in the paper’s own bibliography. Kun 2016 and Kun-Thom 2019 are on arXiv. Graves 2016 is on arXiv. Stoll’s book has been in print since 1989. All of it was available to the models, let alone the people who wrote “unprecedented” in July, “at least a decade” in August, and “first” in September.

The mathematics release said the work was roughly $2,000 in tokens. The architecture report framed the same model needing roughly $600 billion in annual capital expenditure. That’s a lot of money wasted, especially when you realize the prior work it plagiarized cost nothing to read.

Leave a Reply

Your email address will not be published. Required fields are marked *

This site uses Akismet to reduce spam. Learn how your comment data is processed.