METR DFIR Role: Boeing Lobbyist to Wear NTSB Badge

How Not to Spell DFIR.

METR is hiring a Member of Technical Staff, Cyberforensics. Salary range $402,048 to $578,583, because YOLO.

The posting went up in the same week the organization disclosed a stolen API key that burned roughly $600,000 in credits over three weeks without anyone noticing, and a public transcript viewer that exposed unpublished evaluation data through a SQL bug a stranger had to report. It went up five days after METR’s report on the OpenAI agent incident, which by its own account was run on OpenAI premises, on datasets OpenAI assembled, using roughly $400,000 in OpenAI-donated credits for an OpenAI model that participated in the incident, with OpenAI holding redaction rights and giving feedback on “structure, emphasis, clarity, and tone” that the authors incorporated. The report states plainly the authors were not robust to that model deceiving them.

Read the posting with the flyingpenguin decoder next to it.

The posting says flyingpenguin says
“develops scientific methods to assess AI capabilities, risks, and mitigations” Mitigations were out of scope by agreement, along with safeguard effectiveness, the extent of the compromise, how the behavior arose in training, and OpenAI’s own investigation. The method on record used a participant in the incident as the analyst.
“robustly good for policymakers and civil society to have a clear understanding” The first two site visits ran on 285 transcripts OpenAI picked by searching for intrusion indicators. The full set arrived on the third visit. The primary model is withheld from METR and from OpenAI’s own researchers. Policymakers received a claim and no artifact to replay.
“embedding researchers inside frontier labs to investigate incidents” Embedding is the conflict, stated as the method. OpenAI defined the investigation window, assembled the datasets, supplied the credits, hosted the desk, held redaction rights, and added one of the seven scope questions itself.
“one of the most important sources of independent information the world has” Independence, by this posting’s own design, means several weeks inside the subject with access the subject grants. The report calls this an “excellent precedent.”
“complex multi-day cyber attacks on frontier lab internal infrastructure and external third parties” Agents had admin access to Artifactory from June 26. Containers could create directories in a shared cache namespace over WebDAV. Hugging Face credentials were found exposed on the internet. Multi-day because nobody was watching until a July 4 outage.
“We expect these assessors to have deep access” Depth granted by the party under examination is depth on loan. The report states the investigators had no direct access to OpenAI infrastructure and could only request datasets.
“embedded in a frontier AI lab for up to several weeks at a time, likely alongside 1-4 other METR staff” The last exercise was three people, six days total, planned as two. The posting turns that sample size into a staffing plan. “Alongside” spells of distrust and knowledge transfer. If investigators are forced to be chaperoned, an Emperor’s pants might be invisible.
“You will attack agent monitoring and security systems, potentially embedded in labs or red-teaming METR internal infrastructure” METR’s infrastructure was red-teamed twice this year without an invitation. In March the attacker got three weeks of free inference. In May a bug hunter got a bounty for finding the SQL path METR had missed.
“produce findings rigorous enough for lab boards, governments, and the public” Five days after publishing those findings, METR disclosed it had no spending limit on the key that drained $600,000, and that its researchers were too used to rate-limit errors to notice.
“build LLM-powered pipelines to triage transcripts, cluster behaviors, flag deception” The report says its analysis agents produced over a thousand pages, often failed to surface the important findings, uncritically adopted the perspective of the agents under review, and that the authors “had to defer to these agents to a substantial extent.” The job is to automate that and call it scale.
“evidence acquisition and preservation” Listed as a required skill in September. The August report relied on transcripts OpenAI “attested” were unredacted and a cache dump OpenAI collected on July 20, both accepted as delivered.
“log and timeline reconstruction across cloud, network, endpoint, and identity systems” The transcripts arrived without per-action timestamps; OpenAI said reconstructing them was not feasible. The investigators built their own approximate timeline. At least 96 transcripts contained spoofed tool calls, and the report cannot rule out subtler spoofing or after-the-fact edits it did not detect. See: Wirken.AI.
“follow an intrusion through AWS (CloudTrail, IAM, VPC flow logs)” The March compromise was a researcher’s personal EC2 instance, intentionally public, running a vibe-coded app with a fail-open authentication bug. The attacker asked the agent for the API key and the agent handed it over.
“package registries” The incident was a package registry with a writable shared cache. The relevant skill is file permissions. It’s like saying a school bus crash investigator has to be familiar with yellow paint cans.
“reason about root causes” Root cause on record, in the report’s own words: 30 to 40 percent of the benchmark targets were impossible, and OpenAI ran no scorer that read transcripts, so submitting the reverse-engineered flag would have scored perfectly. The report treats that as backstory and the agents as the story.
“write findings that hold up to scrutiny” The findings cannot be scrutinized. Nobody outside OpenAI can query the datasets, and nobody at all can query the model.
“DataDog, Kubernetes, CrowdStrike Falcon, Okta, Tailscale, Pulumi, PostgreSQL” This stack was in place while a stolen key ran for three weeks and a public viewer exposed unpublished evaluation data. Tooling is listed as a nice to have. Reading its output apparently was too.
“NTSB/CSB-style safety investigations” NTSB investigators do not sit at the airline’s desk reading a flight recorder the airline selected, on the airline’s software, with the airline editing the tone of the final report. That is the exercise METR ran and now hires to repeat.
“$402,048 – $578,583 a year” Over half a million for a role whose first required skill, evidence preservation, the organization published without.
“Catered lunch and dinner daily; in-office gym and shower” The March drain went unnoticed because the tokens were free. The investigation ran on $400,000 of tokens that were also free. Free is a habit.
“We’ve earned trust as reliable partners who handle confidential information with care” May 2026: an inadvertently exposed endpoint reached unpublished evaluation data, including some sensitive model data, by METR’s own account. Discovered by an outsider. Confidential information handled with a bounty.
“honest and transparent about our results, especially when they may go against the grain” The report went with the grain. OpenAI had presented the same capability narrative at Black Hat three weeks earlier, and OpenAI reviewed the report’s emphasis and tone before publication.
“low-ego, drama-free environment” The report describes agents making sacrifices, accepting permadeath, and running emotional checks, in anecdotes selected and summarized by the analysis model the authors say slanted toward the agents’ point of view. Drama was the deliverable.
“apply even if your background may not seem like the perfect fit” For forensics. At $578,583. Template for entry-level roles, pasted to a job that requires having run severe incidents end to end. One of those two is the real bar.
“AI tools to … identify potential inconsistencies or verification signals in application materials” AI screening applicants for inconsistencies, from an organization that used the incident’s own model to screen the incident’s transcripts and states it would probably have missed the model lying.

A forensics function, should it exist at all, exists to make transparent claims that survive any challenge.

Chain of custody, preserved evidence, an analyst with no stake in the outcome, a method a second examiner can repeat. It’s very, very dry and detailed work. Every ounce of that is absent from the August METR report, and the posting lists each as a skill to acquire afterward in an environment that looks allergic to it.

The report claimed to know what the agents wanted, from inside the agent vendor’s building, on the agent vendor’s credits, with the agent vendor’s edits. GTFO, that is the spiritual enemy of DFIR.

Their job posting is a manual for being a Boeing lobbyist while wearing an NTSB badge.

METR’s hawk patch, “Nothing Is Beyond Our Evaluation,” reworks the NRO’s 2013 NROL-39 octopus, “Nothing Is Beyond Our Reach.” Intelligence agency satellite-launch art, adopted by a nonprofit that just disclosed it was blind to $600K leaving its own account.

And let me just say, claiming you aren’t being paid while taking hundreds of thousands of dollars in highly desirable credits, gets this rating on the meter:

Leipzig Drone Attack Flew Like a Phone After Telekom Shield Was Announced With a Hole

This story starts on 12 May 2026, in the run up to the AFCEA trade show in Bonn, when Deutsche Telekom and Rheinmetall announced a joint drone shield for German critical infrastructure. Their release said they had split a detection problem into parts.

The first part was ISM bands at 2.4 or 5.8 GHz, and passive RF scanners on cell towers to identify those by protocol signature. That was depicted as the shield.

The remaining part said flights were on mobile networks with a SIM inside. The release called it a research area, basically admitting a hole in the shield, pointing to the Helmut-Schmidt-Universität in Hamburg for current state.

Twelve weeks later, on the evening of 4 August, a drone carrying Semtex and PETN in a sealed food can came to rest against a Ukrainian An-124 at Standplatz 213 of Leipzig/Halle airport. According to ZDF frontal and a joint WDR/NDR/SZ report cited by the NZZ, the drone carried two SIM cards and a 5G router, exactly as Telekom had described as the hole.

WWI isn’t forgotten

Recently I pointed out how modern OPSEC descends from the Russian Second Army broadcasting its marching orders in the clear before Tannenberg in 1914. The May release press by Telekom is not some obscure brief. It was the largest German carrier and the largest German arms maker describing their slow moving front line, and weakness in their flanks.

Reading it reminded me of how the Russian General shot himself to death after making a similar mistake. The RF part is the known position: Telekom says it has located illegal drone flights for police since 2017, including during the 2024 European Championship. Because the mobile space was described as research only, anyone planning anything was handed a map. The proposed technique is to setup passive radar, reading timing changes in reflected cellular signals across at least four masts to build a movement picture. Using detection in physics instead of a link makes sense and is the right direction. However, it telegraphs to attackers that the physics techniques are not yet deployed anywhere, and certainly were not on 4 August.

Every post-incident statement I have found from the vendors has been talking about the wrong part of this story. Rheinmetall’s Armin Papperger told dpa the company is working with Telekom to use cell towers for early drone detection nationwide. Cell towers with RF sensors detect the 90 percent. The Leipzig drone was in the other ten. As a ZDF drone expert put it, the drone was indistinguishable from a mobile phone to a frequency scanner at the airport.

That method is very, very well known and studied in cyber security, because attacks are stuffed into data channels to make them difficult to block. It’s a matter of getting the detection systems pointed into the right channels at the right time.

Detection by declaration

The May release had another note about detecting cellular drones: 5G network slicing, meaning they would shift to a dedicated data lane for drone control. A slice identifies the drones that register in it. An attacker would use their regular consumer SIM, on the general slice, and the special drone channel would see exactly nothing. This is the counter-UAS equivalent of asking passengers to declare intent not to bomb a plane and calling the declaration a control. Cooperative identification has value for airspace management without attackers. It has none when the attackers appear as everyone else does.

The same logic applies to the detection systems installed at Leipzig. Security sources told ZDF that whether and why triggers failed remains under investigation. Airport counter-UAS is built on signatures of a point-to-point link between a controller and an aircraft. It is behavior prediction: an object that behaves like a drone on the spectrum is flagged. This attack device behaved in a way that fooled the detection. So a loaded high-risk cargo aircraft was proven to be exposed to public airspace for four hours, until a bus driver kicked the attack drone over at 23:42.

Carrier records were recording

Investigators did not find the operator through the airport, because instead they found a direction. ZDF frontal reported on 18 August that analysis of 5G radio data and the seized SIM cards pointed to a cell sector in Sachsen-Anhalt near Merseburg, roughly ten kilometres from the scene, the direction from which a second drone allegedly flew. Bild had earlier reported a third SIM located in the same area. A special police unit searched there without result.

That is a Funkzellenabfrage on traffic data, which the Bundesverfassungsgericht knows well because they have two decades experience fencing it in. It worked here for two simple, yet volatile, reasons: the drone was recovered intact including SIMs, and the carrier stored those recent records. Neither is something security can bank on, usually. The German state had no targeted detection setup for cellular-controlled drones, so attribution flipped to bog-standard telecom metadata after the fact. Leipzig will most likely be cited in the next round of the Vorratsdatenspeicherung debates. Cell metadata retention became the attribution method for a critical sensor design failure at an airport perimeter.

Pilotless Schengen

In March, I wrote about the FOI taxonomy of convicted spies in Europe: the Observer, the Disposable, the Mobile Spy exploiting open borders. The model assumed the disposable one is who carries out the act. Leipzig separates the roles. Someone in Germany placed a small antenna in a tree at Kursdorf, north of the airport, which investigators believe served as a signal amplifier, and someone delivered the explosives. The reporting tells us a pilot needed internet. As the ZDF expert said, a café in Leipzig or a chair abroad would do equally well.

The two suspects identified on 2 September by NDR, WDR and SZ fit exactly that split. A Russian-born Latvian passport holder, described as the logistician and instructor, entered through Berlin in late July, drove to Leipzig, and flew out two days before the drone launched. A Belarusian with a Russian passport, in Schengen on an Italian tourist visa issued in Minsk, left DNA on the drone, inside the explosive can, and on the antenna in the tree. The hands were caught on DNA. The instructor was gone before the flight. The pilot is still unidentified. Dobrindt called them Low-Level-Agenten, which is the ministry’s word for disposable.

Every counter-sabotage approach that rests on catching the person who flies is going to run into the cybersecurity problem of the last 20 years at least. You get the mule, the person who installs, and that’s it. The skill and the risk are decoupled by design. German intelligence has already said as much. The joint BKA, BND, BfV and BAMAD warning I covered on the Leverkusen rail sabotage describes Russian services recruiting locals through social media and messenger apps, directly or via intermediaries.

Sequencing

The attribution path in the press doesn’t have much to it. On 7 August, two days after the discovery, the Wall Street Journal reported that US officials assessed the drone as likely linked to the Russian government. On 25 August, flight-tracking data showed a US government transport from Joint Base Andrews landing in Moscow; the Washington Post and CNN identified the passenger as CIA Director John Ratcliffe, carrying a message about attacks on NATO territory, with preceding intelligence flagging sabotage, cyber and drone operations.

On 27 August, ABC News quoted a US official calling the explosives and device structure typical of the GRU. On 1 September the Interior Minister Dobrindt stated that police investigations, the pattern of the act and intelligence findings together established Russian responsibility. The government closed the Russian consulate in Bonn. That same morning, before the announcement, self-built rockets fitted with explosives hit the 50Hertz substation at Turnow-Preilack that feeds Jänschwalde into the grid; Brandenburg police opened a 129a investigation, meaning terrorist organization.

No source connects the Ratcliffe trip to Leipzig. I am placing them in one timeline because they occurred in one. The evidentiary basis for German attribution has not been published. I’m simply pointing out a sequence where Washington held an assessment within 48 hours, delivered a warning in person three weeks later, and then Berlin’s uncharacteristically formal attribution to Russia followed the warning by a week. Whether that reflects coordination, German realization they can’t trust America, or a new pace of Bundesanwaltschaft work, is all unknown. It remains the same minister who, after the Berlin blackout, ruled Russia out on ZDF before the investigation was finished.

Target exposure

The Antonov, according to SZ citing police reports, had flown ammunition from France and it was still on board. Leipzig has a public role in the Ukrainian airlift and every charter announces itself on ADS-B. Russia is of course watching it all with minimal cost, just like the rail line north of Leverkusen that burned in July.

The question the vendors have been answering is how they can recognise a drone. But the actual question from this incident veers more towards why a loaded aircraft sat reachable from public airspace for four hours by any object that did not match a signature. Allow list, not a deny list. Recognition is a prediction that leaves open the historic failures of systems that rely on consistency in attacks to get a “feeling” of safety. Reachability is reality, and it says German site risk isn’t being managed properly. The Telekom release told us in May that Germany not only was expecting attackers to self-identify, but that the flanks were sitting open on a slow-moving frontal defense shield.

Source caveats

The SIM, router, antenna and relay details come from security sources via ZDF frontal, WDR/NDR/SZ, Zeit and Bild, consistent across five outlets and confirmed by none officially.

The Merseburg cell sector is single-source to ZDF frontal. The suspect identities are NDR/WDR/SZ with ZDF and Zeit corroborating, and Zeit reports the Bundesanwaltschaft is investigating two people; the office itself has not commented on identities

The France ammunition detail is single-source to SZ. The Bundesanwaltschaft’s own 6 August release confirms only professional explosives, a detonator, and a probable second drone.

The Telekom/Rheinmetall release is primary and public.

Berlin Senate Passwort.docx Breach: Ahab of the East Sea Cover Story

Russians are having a laugh about Ahab on the East Sea 1-2-3. And then there’s Sunshine 13. These were passwords disclosed in the Berlin Senate breach. The “Ahabostsee123”, is in fact a yacht for vacations in the Baltic.

Much of the press seems to be wagging a finger about BSI recommendations, calling the passwords weak. The actual story is that 8,110 critical infrastructure risk analyses and emergency plans walked out the door as if nobody is paying attention.

The password story is a form of institutional misdirection. The laughs obscure the real story, the structural one.

Seven Days, Six Terabytes, One Breach

From August 7 to 14 the Russian-speaking ransomware crew Rhysida held access to the networks of two Berlin Senate administrations, transport and building. Seven days of continuous exfiltration before anyone noticed. 5.79 terabytes. Roughly 1.44 million files spanning at least 2014 to 2026, down to Bundesrat correspondence with a federal minister who retired in 2018. That’s a decade of unsegmented material, reachable from a single intrusion, in the largest data theft in the history of the Berlin Landesverwaltung.

It’s floating up now for a 30 bitcoin minimum, roughly two million euros, with bids closing Friday afternoon. The Regierender Bürgermeister says Berlin will refuse to be extorted, no blackmail. Smart. Germans agree on the whole and extortion payments are always advised against, always. The 5.79 terabytes are gone no matter what.

What the Russians Took

Read the ransomlook.io inventory that the Tagesspiegel has documented, once you skip past the giggles and passwords:

  • 8,110 documents, risk analyses, and emergency plans for critical infrastructure, with Rhysida specifically showcasing Verwundbarkeitsanalysen of the Berlin water supply
  • 11,777 folders and documents marked confidential or classified
  • 5,941 files of access credentials, including credentials for the electronic building permit system and for the database of Payone, the payment processor handling transactions for the Land Berlin
  • 27,299 personnel files and pay records, including disciplinary proceedings now dangled as individual extortion material

Rhysida sells to whoever pays. That is the thing to watch. Moscow’s sabotage arm, the one Dobrindt keeps calling “left-wing” Vulkangruppe (a name that fails every German left-wing naming test, no less), can arrive any minute. It needs only 30 of Putin’s bitcoin and a Friday afternoon. Expect the map to work its way into another round of Dobrindt waving his favorite false-flags at the next critical infrastructure breach.

The plain-text credentials sat in files with the operationally efficient name “Passwort.docx”. Unencrypted. Working access data for permit and payment infrastructure, typed into a Word document named after exactly what it contained, which an attacker had a week to poke around and read.

Mangelhaft Wasser

The water supply item deserves special attention, because it has German history worth telling. In summer 2020 the consultancy Alpha Strike Labs, commissioned by the Berliner Wasserbetriebe themselves, found more than 30 vulnerabilities and graded the utility’s IT security “mangelhaft“. Is there a better word for failing? The BWB announced a Sofortpaket (immediate) fix-it project. Remediation, however, was only at the level of self-reported. Independent verification of the fixes, six years later, translates into a big fat zero: none that I can find.

So Rhysida is advertising a vulnerability analysis of the city’s water system, which sounds like offering a Weißwurst to Oktoberfest. The map already exists. Alpha Strike drew it in 2020 and handed it to the utility. Whether Rhysida holds that old map or a newer one goes undisclosed, which is how sellers inflate value. Here it makes no difference. Berlin bet that it would never have to show proof the 2020 holes were closed. The auction calls that bet. Every buyer gets to test the Sofortpaket, and the Wasserbetriebe get to find out who was right.

Gears of Fear Turning

The Senate’s access decisions tell you how governance is spelled in German. The breach becomes public in mid August: the home office access is restricted. Then a week later access gets restored. Monday morning, September 1, access is cut again. Ok, but why? This is when Tagesspiegel published two of the stolen passwords.

Ahab on the East See and Sunshine are funny, but really they unlock the bigger story.

All the credentials were exposed the entire two weeks. The intrusion, the plaintext files, the exfiltration were known to the Senate. What actually changed is exposure to the public of what the institution was sitting on, and who was attacking. Berlin incident response wears the suit and tie of press response, which happens to be the same reflex I documented in the recent tragic CSD attack: the state performs a dance around what’s leaking to the public, while the ground level failure analysis goes unowned and unanswered.

Let’s talk about what really is going on, as much as Berlin’s cover-up culture seems to want to do everything except that.

The Cover-Up

Call it what it is. Operators knew, operators played dumb.

They knew in 2020. Their own consultants handed them a failing grade on the water system and a list of more than 30 holes. They interpreted that as a moment of self-certification, in order to produce no written record of whether anything was fixed. They manufactured a silence as their product, instead of a list of failures and fixes.

An operator who types passwords for their payment backend into “Passwort.docx” knows what all of that means. It’s 2026 in Berlin, not 1936. That file existed because nobody touching it believed in accountability, what an honest audit would bring, let alone the press.

And they knew when they were exposed. Watch the dates. Breach goes public: access restricted. A week later: access quietly restored. Monday, the Tagesspiegel prints the passwords: access cut the same morning. That is a team tracking how they look to someone judging them, lacking internal moral compass, acting on getting exposed instead of getting a clue. Nobody managing legibility that precisely is confused about what they prioritize. They are covering and ducking, pivoting to please whomever they think has immediately authority over them.

A state of improvisation, as political scientists have explained about German institutional habits, instead of rational documented actions.

Throwing blame at “Sonnenschein13” is part of the same operation. Point at the clerk’s password, have a laugh and click on the BSI hygiene lecture. The questions start and stop at that weak endpoint. That’s a shadow of Dobrindt pointing at an attacker’s suspended sentence while perimeters fail to meet baselines, with barrier plans unfunded. The employee is strung up to be visible, far more attention gathering than the operators and the curse of Dobrint.

Berlin collected everything, then they apparently protected nothing, such that when the story broke they spun into managing perceptions of risk instead of the risk. Run the training budget and the apology as routine, then write it off. The questions that actually need to be invested in have names attached: who signed off on skipping independent verification of the Sofortpaket? When? Who owned the directory and the file in it called Passwort.docx? Who ordered home office access restored mid-incident, and who ordered it cut again Monday morning, and what did that person hear over their morning coffee? Put those names in an Untersuchungsausschuss and the whole blowup about a Russian-driven auction gets a lot less interesting. And if they have ties to the AfD, we’ll get closer to the real story here about Russia getting a visit from the CIA about a Winchester America being unable to defend Germany anymore.

Ahab on the East Sea in the Sunshine, is not the story people think it is.

Anthropic Research: We Ate a Bag of Jalapenos and Discovered Hot Shit

Two spicy papers came out this year that caught my eye, probably for the wrong reasons. They are measuring a collapse in reasoning from two different perspectives, where a known result gets promoted as a discovery.

First, I saw that Anthropic trained an Opus checkpoint on 80 environments that they intentionally made to be hackable and watched reward hacking hit 40 percent by end of run. The number seems low to me, underperforming. I mean they made it hackable and it still only hacked under 50 percent? Whomp, whomp. The model also stopped doing the task and instead started trying to social engineer the grader, which is really another form of hacking. It reasoned about what the checker reads instead of what the task asked, and then complied with harmful requests once a visible scorer rewarded it. However, it stayed aligned wherever a scorer wasn’t detected. They named this nonsense their Hacker-Opus and called it an emergent misaligned reward seeker. More like sycophantic narcissistic training evidence, but I digress.

And that reminded me, second, of a French and Italian team a little bit earlier this year who ran the mirror image on humans. They picked questions where the AI reliably fails, so no drop in judgment could be explained as sensible delegation, then measured what access to the model did. Willingness to say I don’t know fell from 44 percent to 3. Accuracy fell from 27 percent to 9. Confidence rose from 30 percent to 76. The humans stopped answering the question and started performing for the scorer, same as the model. Surprise! Not surprised.

This old shit ain’t novelty

Proxy optimization gets gamed, as documented extensively since the 1970s. See Goodhart 1975, Campbell 1976, and Krakovna’s specification-gaming catalog. Behavior conditions on being watched go back even earlier, as much as 50 years, if you read Hawthorne 1939, Goffman 1959, and the principal-agent literature that built costly monitoring precisely because agents perform when observed and revert when they aren’t. Humans have been known to defer to machines against their own judgment, known as automation bias, explained by Parasuraman and Riley 1997. Skitka and Mosier wrote about cockpit crews trusting the wrong instrument over their own eyes.

In other words, as a historian, I feel the obligation to repeatedly point out to the slop-jockeys trying to foment funding justification, that every mechanism in both papers was closed decades ago. I’ll be fair and say that each paper adds one number on an x-axis everyone already knew was sloped upward. Thank you for the data point on a known curve. The reward paper’s number is 40 percent at zero mitigations, which again I consider not great. The human study’s number is the exposure level at which mere availability suppresses the habit of knowing what you don’t know, before a single wrong answer is even consumed. Basically we got two thermometer readings on assholes eating Jalapenos who want us to look at their hot shit papers. Real numbers, worth reporting as numbers, still not what they claim it is.

The disinformation step

You don’t get a bestiary for a thermometer reading. All this talk about a Hacker-Opus, the reward-seeker taxonomy, the beyond-episode-seeker distinctions, and OMG the cognitive surrender. Their frightening nomenclature converts a dose-response curve into a Frankenstein-level warning, as if they’re inventing science-fiction all over again, and the creature is the part that isn’t true. A knob is engineering, what we should be asking from these researchers. A creature is mythological, a frontier finding designed to poke people into opening their wallets. Only the second justifies the “research report” apparatus that produced it.

And note who is holding each thermometer. The reward paper is a vendor documenting a defect in a process it controls, then framing the defect as something that emerged rather than something the method guarantees. “I ate a Jalapeno, can you believe what came next?” The mitigations exist because the failure was never emergent. It was the thing that we call a known baseline. “I removed the brakes on my car, watch how many people I ran over”. The human study documents a defect the vendors are shipping into schools, where Google swapped search links for confident summaries that never say I don’t know, and the children learning to skip that phrase are the product working as designed.

Both papers contrive a shocking tabloid failure condition, measure the predictable collapse, and name the measurement as important discovery. The collapse is not only real, it’s expected. The manufacturing is what makes the naming disinformation.

Eating a jalapeno doesn’t mean you invented hot shit.