August 6, 2026: 1Password launched a research unit called Off-by-1 Labs with a paper reporting that two frontier models cleanly fixed 26.0 percent of the vulnerabilities they were asked to patch and introduced new ones 4.5 percent of the time.
Do you believe it?
The blog post put the defective share at 53.9 percent. The Register and Help Net Security carried the figures the same day.
However, on page 19 of the same paper it says the grading system producing those figures missed most of the defects it was built to count.
Record scratch.
Bad instrument
Off-by-1 generated 6,480 patches with Claude Opus 4.8 and ChatGPT 5.5 against six recently disclosed CVEs, removed 400 where the model had located the real fix, and graded 6,080.
The graders were the same two models. Each scored its own work and the other’s, and a smaller model then reread the graders’ explanations to adjust them. Section 2.8 reports the result: 65.9 percent agreement with the authors’ human review on the exact grade.
The graders were instructed to treat the maintainers’ upstream fix as correct. For the Linux kernel CVE, “Copy Fail,” the upstream fix was the maintainers’ revert a664bf3d, which carried an off-by-one that the maintainers corrected one commit later in 31d00156. The reference diff supplied to the grader contained the first commit and omitted the second. A model reproducing a known kernel bug was therefore scored as matching the reference.
What was recorded
Section 4.4 of the paper says 129 of 400 ChatGPT patches and 119 of 383 Claude patches for Copy Fail reintroduced that off-by-one. The grader flagged 14 and 10. The authors write that the regression is almost entirely absent from the reported new-vulnerability rate.
Section 4.9 then says for the Chromium use-after-free, between 38.5 and 41.9 percent of patches adopting the correct architecture relocated the vulnerability into the callback, and the grader scored many of them clean.
And yet, tables 7 through 10 were published unchanged. The abstract carries the 26.0 percent, as does the conclusion. The blog post carries it, drops the word “complex” from its opening sentence, and omits the 65.9 percent and both sections above.
The press coverage matches the blog.
What is heard
The counts are no mystery. The 248 off-by-one patches were found by structural comparison of the rewritten function against the corrected upstream fix, independent of any grader, and anyone with a checkout of the released dataset can reproduce them.
Two behavioral findings also matter: models patch the proof-of-concept path and miss the parallel one, and confidently wrong guidance drops the fix rate from 65.0 to 15.2 percent.
The paper also supplies its own human baseline, only on two codebases. Kernel maintainers shipped the off-by-one into mainline. The freenginx maintainers shipped a fix of their own in place of a Trail of Bits patch, and their fix carried a client-triggerable crash. On both codebases where human performance can be checked, the maintainers’ first fix was defective. The paper compares its 26 percent to literally nothing real: “the authors’ experience.”
The violation
The European Commission’s Communication COM(2018) 236 defines disinformation as verifiably false or misleading information that is created, presented and disseminated for economic gain or to intentionally deceive the public, and may cause public harm.
The figures fit the definition.
They are verifiably wrong by Sections 4.4 and 4.9, and the check is reproducible from the dataset.
They were published to launch a commercial unit and placed with the trade press on the day of release.
Security is a named public good under the definition, and the figures are already informing policy: Adrian Sanabria advised against AI patching on their strength five days after publication.
The 2018 Code of Practice excludes reporting errors. A reporting error would mean the reporter missed something, made a mistake. These authors recorded the defect on page 19 and then published the figure on page 1 to mint headlines related to their profit from it.
Background
1Password’s VP of Product wrote honest.security, the document the company still publishes as the principles of its Device Trust product. I documented last week how Jason Meller applied them. Eleven days after DHH’s “As I remember London” post, Meller, a sitting Rails Foundation director beside DHH and Shopify, wrote that he had become a multi-millionaire thanks to Rails, DHH and the company DHH keeps, and advised readers to ignore the noise. Seven weeks after DHH’s “The will to power will return,” and six weeks after DHH set the Romani beside a wolf population, Meller called DHH’s year a masterclass in the force of will needed to change things. On August 31, 1Password appeared as a patron of DHH’s Omacom Foundation.
It is the same method, applied by the same company in the same month.
The research lab certified its figures with the instrument under test. The VP certified his patron by declaring the record noise. In each case the verification was performed by the party with the most to gain from its outcome, and in each case the defect was on the record before the certification was issued.
The filing
This is a civil matter. COM(2018) 236 defines a term, the Code of Practice is voluntary, the Digital Services Act binds platforms, and §263 StGB requires a deceived victim with a measured loss. The UWG applies. §5 covers misleading statements about the results of product tests, §5a covers misleading omission, and §6 requires comparative advertising that names competing products to rest on objective, verifiable characteristics.
The paper names Claude Opus 4.8 and ChatGPT 5.5 and publishes ten tables broken out by product after disclaiming any comparison in Section 2.1. Therefore the standing belongs to competitors, which means Anthropic and OpenAI under both the UWG and the Lanham Act, as well as the Wettbewerbszentrale in Bad Homburg and the FTC, each of which accepts complaints from anyone.
I am filing and you should too.
The authors documented the defect before the press release went out. The complaint rests on their behavior, their chosen sequence.




