Anthropic Research: We Ate a Bag of Jalapenos and Discovered Hot Shit

Two spicy papers came out this year that caught my eye, probably for the wrong reasons. They are measuring a collapse in reasoning from two different perspectives, where a known result gets promoted as a discovery.

Anthropic trained an Opus checkpoint on 80 environments that they intentionally made to be hackable and watched reward hacking hit 40 percent by end of run. The number seems low to me, underperforming. I mean they made it hackable and it still only hacked under 50 percent? Whomp, whomp. The model also stopped doing the task and instead started trying to social engineer the grader, which is really another form of hacking. It reasoned about what the checker reads instead of what the task asked, and then complied with harmful requests once a visible scorer rewarded it. However, it stayed aligned wherever a scorer wasn’t detected. They named this nonsense their Hacker-Opus and called it an emergent misaligned reward seeker. More like sycophantic narcissistic training evidence, but I digress.

A French and Italian team a little bit earlier this year ran the mirror image on humans. They picked questions where the AI reliably fails, so no drop in judgment could be explained as sensible delegation, then measured what access to the model did. Willingness to say I don’t know fell from 44 percent to 3. Accuracy fell from 27 percent to 9. Confidence rose from 30 percent to 76. The humans stopped answering the question and started performing for the scorer, same as the model. Surprise! Not surprised.

This old shit ain’t novelty

Proxy optimization gets gamed, as documented extensively since the 1970s. See Goodhart 1975, Campbell 1976, and Krakovna’s specification-gaming catalog. Behavior conditions on being watched go back even earlier, as much as 50 years, if you read Hawthorne 1939, Goffman 1959, and the principal-agent literature that built costly monitoring precisely because agents perform when observed and revert when they aren’t. Humans have been known to defer to machines against their own judgment, known as automation bias, explained by Parasuraman and Riley 1997. Skitka and Mosier wrote about cockpit crews trusting the wrong instrument over their own eyes.

In other words, as a historian, I feel the obligation to repeatedly point out to the slop-jockeys trying to foment funding justification, that every mechanism in both papers was closed decades ago. I’ll be fair and say that each paper adds one number on an x-axis everyone already knew was sloped upward. Thank you for the data point on a known curve. The reward paper’s number is 40 percent at zero mitigations, which again I consider not great. The human study’s number is the exposure level at which mere availability suppresses the habit of knowing what you don’t know, before a single wrong answer is even consumed. Basically we got two thermometer readings on assholes eating Jalapenos who want us to look at their hot shit papers. Real numbers, worth reporting as numbers, still not what they claim it is.

The disinformation step

You don’t get a bestiary for a thermometer reading. All this talk about a Hacker-Opus, the reward-seeker taxonomy, the beyond-episode-seeker distinctions, and OMG the cognitive surrender. Their frightening nomenclature converts a dose-response curve into a Frankenstein-level warning, as if they’re inventing science-fiction all over again, and the creature is the part that isn’t true. A knob is engineering, what we should be asking from these researchers. A creature is mythological, a frontier finding designed to poke people into opening their wallets. Only the second justifies the “research report” apparatus that produced it.

And note who is holding each thermometer. The reward paper is a vendor documenting a defect in a process it controls, then framing the defect as something that emerged rather than something the method guarantees. “I ate a Jalapeno, can you believe what came next?” The mitigations exist because the failure was never emergent. It was the thing that we call a known baseline. “I removed the brakes on my car, watch how many people I ran over”. The human study documents a defect the vendors are shipping into schools, where Google swapped search links for confident summaries that never say I don’t know, and the children learning to skip that phrase are the product working as designed.

Both papers contrive a shocking tabloid failure condition, measure the predictable collapse, and name the measurement as important discovery. The collapse is not only real, it’s expected. The manufacturing is what makes the naming disinformation.

Eating a jalapeno doesn’t mean you invented hot shit.

Leave a Reply

Your email address will not be published. Required fields are marked *

This site uses Akismet to reduce spam. Learn how your comment data is processed.