Category Archives: Food

Anthropic Research: We Ate a Bag of Jalapenos and Discovered Hot Shit

Two spicy papers came out this year that caught my eye, probably for the wrong reasons. They are measuring a collapse in reasoning from two different perspectives, where a known result gets promoted as a discovery.

First, I saw that Anthropic trained an Opus checkpoint on 80 environments that they intentionally made to be hackable and watched reward hacking hit 40 percent by end of run. The number seems low to me, underperforming. I mean they made it hackable and it still only hacked under 50 percent? Whomp, whomp. The model also stopped doing the task and instead started trying to social engineer the grader, which is really another form of hacking. It reasoned about what the checker reads instead of what the task asked, and then complied with harmful requests once a visible scorer rewarded it. However, it stayed aligned wherever a scorer wasn’t detected. They named this nonsense their Hacker-Opus and called it an emergent misaligned reward seeker. More like sycophantic narcissistic training evidence, but I digress.

And that reminded me, second, of a French and Italian team a little bit earlier this year who ran the mirror image on humans. They picked questions where the AI reliably fails, so no drop in judgment could be explained as sensible delegation, then measured what access to the model did. Willingness to say I don’t know fell from 44 percent to 3. Accuracy fell from 27 percent to 9. Confidence rose from 30 percent to 76. The humans stopped answering the question and started performing for the scorer, same as the model. Surprise! Not surprised.

This old shit ain’t novelty

Proxy optimization gets gamed, as documented extensively since the 1970s. See Goodhart 1975, Campbell 1976, and Krakovna’s specification-gaming catalog. Behavior conditions on being watched go back even earlier, as much as 50 years, if you read Hawthorne 1939, Goffman 1959, and the principal-agent literature that built costly monitoring precisely because agents perform when observed and revert when they aren’t. Humans have been known to defer to machines against their own judgment, known as automation bias, explained by Parasuraman and Riley 1997. Skitka and Mosier wrote about cockpit crews trusting the wrong instrument over their own eyes.

In other words, as a historian, I feel the obligation to repeatedly point out to the slop-jockeys trying to foment funding justification, that every mechanism in both papers was closed decades ago. I’ll be fair and say that each paper adds one number on an x-axis everyone already knew was sloped upward. Thank you for the data point on a known curve. The reward paper’s number is 40 percent at zero mitigations, which again I consider not great. The human study’s number is the exposure level at which mere availability suppresses the habit of knowing what you don’t know, before a single wrong answer is even consumed. Basically we got two thermometer readings on assholes eating Jalapenos who want us to look at their hot shit papers. Real numbers, worth reporting as numbers, still not what they claim it is.

The disinformation step

You don’t get a bestiary for a thermometer reading. All this talk about a Hacker-Opus, the reward-seeker taxonomy, the beyond-episode-seeker distinctions, and OMG the cognitive surrender. Their frightening nomenclature converts a dose-response curve into a Frankenstein-level warning, as if they’re inventing science-fiction all over again, and the creature is the part that isn’t true. A knob is engineering, what we should be asking from these researchers. A creature is mythological, a frontier finding designed to poke people into opening their wallets. Only the second justifies the “research report” apparatus that produced it.

And note who is holding each thermometer. The reward paper is a vendor documenting a defect in a process it controls, then framing the defect as something that emerged rather than something the method guarantees. “I ate a Jalapeno, can you believe what came next?” The mitigations exist because the failure was never emergent. It was the thing that we call a known baseline. “I removed the brakes on my car, watch how many people I ran over”. The human study documents a defect the vendors are shipping into schools, where Google swapped search links for confident summaries that never say I don’t know, and the children learning to skip that phrase are the product working as designed.

Both papers contrive a shocking tabloid failure condition, measure the predictable collapse, and name the measurement as important discovery. The collapse is not only real, it’s expected. The manufacturing is what makes the naming disinformation.

Eating a jalapeno doesn’t mean you invented hot shit.

There’s a parasite loose in the salad

There’s a parasite loose in the salad,
And the answer from Trump has been pallid.
In forty-seven states
His agency waits,
Strumming one shitty fast-food ballad.

They made the surveillance elective,
Then said, when the count turned defective,
That FoodNet was never
For outbreaks. How clever.
Week seven. A leisurely detective.


Related: It feels like some people are just discovering the down side to disclosure labels being NOT A CONTROL.

“…making more things does not make me make better things.” And he said that he still needs to come to terms “with the fact that the level of dopamine I’ve been getting from interacting with LLMs… with doing more and more and more and more… is not healthy for me or good for the world.”

But being unhealthy and bad is such a fundamental part of American wealth transfer. What’s next, he lets his staff form a union or they all take a real holiday?

Why Fascists Hate Pasta

Here’s a compelling look at the politics of any dining table. Fascists engineer non-dependence for themselves, and dependence for others. The first enables their aggression and the second finances and shields it. Pasta was the propaganda surface of Mussolini’s war economy.

Marinetti published the Manifesto of Futurist Cooking in the Gazzetta del Popolo on 28 December 1930, calling pasta an “absurd Italian gastronomic religion” that induced sloth, pessimism, and unfitness for war.

Mussolini gave it a sympathetic hearing for autarkic reasons rather than culinary ones: Italy imported wheat, the 1925 Battle for Grain was faltering, and rice from the Po Valley was domestic.

The regime promoted rice through the Ente Nazionale Risi, but pasta was never banned, and the press caught Marinetti himself eating spaghetti at Biffi in Milan. Then the army was caught too.

The Regio Esercito ration kept pasta, and in North Africa 1940 to 1943 the water required to boil it became a genuine logistical liability, noted in British intelligence assessments comparing Italian and German water consumption in the desert.

The fascist avant-garde declared war on spaghetti, and spaghetti won, as evidenced by the inept North Africa campaign logistics and failed supply columns.

Chickenshit UK Egg Math: 2.67 Million Printed as 2.67 Percent

Somebody really clucked up a new report about eggs in the UK.

The environmental cost of welfare-driven policy changes in UK egg production is a paper that appeared in Royal Society Open Science on 29 July 2026. And then it was scooped up by The Guardian.

Organic eggs have worse impact on climate than eggs from caged chickens

Hold on a second, that’s a sensational claim. Let’s take an actual look at the paper.

It models eight ways of meeting UK egg demand and wants us to believe that furnished cages carry the lowest environmental impact and need the fewest birds.

Should we believe?

I’m here to tell you that in five places its text contradicts its own charts. All five sit in the section on how many hens each system requires. How did this go to print, let alone spur The Guardian to cluck, cluck about what’s “worse”?

Apparently the shift to fully organic free-range production would need 2.67 million additional laying hens. Against a national flock of 40.44 million, that would be an increase of 6.60 percent. The paper, however, writes just 2.67 percent. The difference of millions of chickens has been printed as a percentage. It happens four times in a single paragraph. A fifth figure in the same paragraph is calculated correctly, so it’s some kind of weird sloppy mistake, which everyone can see plainly.

Each of the four wrong percentages is too small by exactly 2.4728. That number is 100 divided by 40.44. So I believe the same simple wrong calculation was repeated four times over.

The vocabulary also matters, given what happens next. “Furnished cage” is the paper’s term. Defra, the government department that regulates this, calls the same system an “enriched colony cage”.

On 12 January 2026 Defra, acting for the UK, Welsh, Scottish and Northern Ireland governments, opened a consultation proposing a ban from 2027 on installing new enriched colony cages and a ban from 2032 on using existing ones. It closed on 9 March. The responses are still being analysed. Enriched colony cages supply a little over a fifth of UK shell egg production.

What the paper prints What its own figures give What happened
Population increase of 1.01, 0.74, 0.79 and 2.67 percent for the four non-cage scenarios 2.50, 1.83, 1.95 and 6.60 percent Absolute millions printed as percentages
Mortality reductions of 0.94 and 1.33 million hens under furnished and battery cage Reductions of 1.51 and 1.12 million Scenario totals printed as differences, reversing which cage system performs better
Eutrophication reduction of 4.89 and 12.20 percent for furnished and battery cage 12.20 and 5.49 percent Transposed, and 4.89 corresponds to nothing in the dataset
Acidification of 41 678 tonnes SO2-eq under cage-free barn 40 678 tonnes Text contradicts the figure; the stated 6.53 percent matches the figure
National flock of 36.72 million under furnished cage 37.66 million Culled subtotal printed as a total, then compared against a battery cage total

Every one of the five drops the paper’s credibility a notch. All together it seems too weird. It wants to argue that cage-free systems need more hens for the same number of eggs. But then three of their five errors weaken that case, while two strengthen it. The mortality error also makes battery cages, banned in the UK since 2012, look better than furnished cages. That’s a clue something is very wrong with their whole setup.

The mistakes do have a pattern. Each wrong figure is a quantity being correctly calculated, and then filed under the wrong heading. 2.67 is the count of extra hens in millions. 0.94 million is total mortality under furnished cages, entered as a saving. 36.72 million is the number of hens slaughtered, entered as the number kept. Every one of them is a bar in Figure 4.

That suggests the pictures were used incorrectly to generate the text. Anyone holding the model divides 2.67 by 40.44 and gets the 6.60. The digits are all there, but they’ve been scrambled.

Supposedly seven authors gave final approval for this math, and agreed to be held accountable. What they approved understates the welfare result the whole argument rests on by a factor of two and a half.

Here are six more things I found while trying to understand why this paper is so bad.

  • The inventory comes from Leinonen 2012 and 2014, which is egg production data from around 2009 and 2010. The paper limitations section does not mention this rather critical point. An author mentions it to the Guardian, however.
  • Mortality is fixed at 2.50 percent in furnished cages rising to 8.00 percent in organic systems, all of it carried over from the same 2009 and 2010 sources. Reference 30 is Schuck-Paim 2021, covering 6040 commercial flocks and 176 million hens in 16 countries, which found no mortality difference between caged and cage-free housing in recent years. That study covers indoor systems only, so it bears on the barn figure rather than the free-range and organic ones. The paper dismisses it in a subordinate clause and cites it with the wrong volume number.
  • Current UK grid emission factors are applied to 2010 energy consumption numbers. That’s wildly out of date. Barn acidification is then attributed to energy use.
  • Welfare is only defined as headcount, mortality and slaughter number, which makes no sense. The paper says openly that quality of life was not available to it. That admission gives up the game. Under a fixed output constraint, a metric that counts animals rewards whichever system converts feed fastest. The finding follows from the setup.
  • The method comes from Hayek and Garrett 2018 and Chan, Franks and Hayek 2022, both cited. Those authors built the headcount frame to argue that extensive systems cannot scale and consumption must fall. This paper places consumption change out of scope, stealing an instrument and removing the important warning labels.
  • Two authors hold commercial affiliations. Two. Margaret Hegwood manages corporate engagement at the Centre for Feed Innovation, a think tank founded in 2024 to advance novel feed ingredients, on a paper finding that feed dominates every impact category. Rhea Harrison is at Foodsteps Ltd, an environmental consulting company that sells life cycle assessments to food businesses and markets its research partnerships as a route to shaping policy. The declaration names Hegwood. That’s only one. And states that the other authors have no competing interest.

The paper page carries a link to its review history. And for this paper, when you click on it, you get this banner instead: