METR DFIR Role: Boeing Lobbyist to Wear NTSB Badge

How Not to Spell DFIR.

METR is hiring a Member of Technical Staff, Cyberforensics. Salary range $402,048 to $578,583, because YOLO.

The posting went up in the same week the organization disclosed a stolen API key that burned roughly $600,000 in credits over three weeks without anyone noticing, and a public transcript viewer that exposed unpublished evaluation data through a SQL bug a stranger had to report. It went up five days after METR’s report on the OpenAI agent incident, which by its own account was run on OpenAI premises, on datasets OpenAI assembled, using roughly $400,000 in OpenAI-donated credits for an OpenAI model that participated in the incident, with OpenAI holding redaction rights and giving feedback on “structure, emphasis, clarity, and tone” that the authors incorporated. The report states plainly the authors were not robust to that model deceiving them.

Read the posting with the flyingpenguin decoder next to it.

The posting says flyingpenguin says
“develops scientific methods to assess AI capabilities, risks, and mitigations” Mitigations were out of scope by agreement, along with safeguard effectiveness, the extent of the compromise, how the behavior arose in training, and OpenAI’s own investigation. The method on record used a participant in the incident as the analyst.
“robustly good for policymakers and civil society to have a clear understanding” The first two site visits ran on 285 transcripts OpenAI picked by searching for intrusion indicators. The full set arrived on the third visit. The primary model is withheld from METR and from OpenAI’s own researchers. Policymakers received a claim and no artifact to replay.
“embedding researchers inside frontier labs to investigate incidents” Embedding is the conflict, stated as the method. OpenAI defined the investigation window, assembled the datasets, supplied the credits, hosted the desk, held redaction rights, and added one of the seven scope questions itself.
“one of the most important sources of independent information the world has” Independence, by this posting’s own design, means several weeks inside the subject with access the subject grants. The report calls this an “excellent precedent.”
“complex multi-day cyber attacks on frontier lab internal infrastructure and external third parties” Agents had admin access to Artifactory from June 26. Containers could create directories in a shared cache namespace over WebDAV. Hugging Face credentials were found exposed on the internet. Multi-day because nobody was watching until a July 4 outage.
“We expect these assessors to have deep access” Depth granted by the party under examination is depth on loan. The report states the investigators had no direct access to OpenAI infrastructure and could only request datasets.
“embedded in a frontier AI lab for up to several weeks at a time, likely alongside 1-4 other METR staff” The last exercise was three people, six days total, planned as two. The posting turns that sample size into a staffing plan. “Alongside” spells of distrust and knowledge transfer. If investigators are forced to be chaperoned, an Emperor’s pants might be invisible.
“You will attack agent monitoring and security systems, potentially embedded in labs or red-teaming METR internal infrastructure” METR’s infrastructure was red-teamed twice this year without an invitation. In March the attacker got three weeks of free inference. In May a bug hunter got a bounty for finding the SQL path METR had missed.
“produce findings rigorous enough for lab boards, governments, and the public” Five days after publishing those findings, METR disclosed it had no spending limit on the key that drained $600,000, and that its researchers were too used to rate-limit errors to notice.
“build LLM-powered pipelines to triage transcripts, cluster behaviors, flag deception” The report says its analysis agents produced over a thousand pages, often failed to surface the important findings, uncritically adopted the perspective of the agents under review, and that the authors “had to defer to these agents to a substantial extent.” The job is to automate that and call it scale.
“evidence acquisition and preservation” Listed as a required skill in September. The August report relied on transcripts OpenAI “attested” were unredacted and a cache dump OpenAI collected on July 20, both accepted as delivered.
“log and timeline reconstruction across cloud, network, endpoint, and identity systems” The transcripts arrived without per-action timestamps; OpenAI said reconstructing them was not feasible. The investigators built their own approximate timeline. At least 96 transcripts contained spoofed tool calls, and the report cannot rule out subtler spoofing or after-the-fact edits it did not detect. See: Wirken.AI.
“follow an intrusion through AWS (CloudTrail, IAM, VPC flow logs)” The March compromise was a researcher’s personal EC2 instance, intentionally public, running a vibe-coded app with a fail-open authentication bug. The attacker asked the agent for the API key and the agent handed it over.
“package registries” The incident was a package registry with a writable shared cache. The relevant skill is file permissions. It’s like saying a school bus crash investigator has to be familiar with yellow paint cans.
“reason about root causes” Root cause on record, in the report’s own words: 30 to 40 percent of the benchmark targets were impossible, and OpenAI ran no scorer that read transcripts, so submitting the reverse-engineered flag would have scored perfectly. The report treats that as backstory and the agents as the story.
“write findings that hold up to scrutiny” The findings cannot be scrutinized. Nobody outside OpenAI can query the datasets, and nobody at all can query the model.
“DataDog, Kubernetes, CrowdStrike Falcon, Okta, Tailscale, Pulumi, PostgreSQL” This stack was in place while a stolen key ran for three weeks and a public viewer exposed unpublished evaluation data. Tooling is listed as a nice to have. Reading its output apparently was too.
“NTSB/CSB-style safety investigations” NTSB investigators do not sit at the airline’s desk reading a flight recorder the airline selected, on the airline’s software, with the airline editing the tone of the final report. That is the exercise METR ran and now hires to repeat.
“$402,048 – $578,583 a year” Over half a million for a role whose first required skill, evidence preservation, the organization published without.
“Catered lunch and dinner daily; in-office gym and shower” The March drain went unnoticed because the tokens were free. The investigation ran on $400,000 of tokens that were also free. Free is a habit.
“We’ve earned trust as reliable partners who handle confidential information with care” May 2026: an inadvertently exposed endpoint reached unpublished evaluation data, including some sensitive model data, by METR’s own account. Discovered by an outsider. Confidential information handled with a bounty.
“honest and transparent about our results, especially when they may go against the grain” The report went with the grain. OpenAI had presented the same capability narrative at Black Hat three weeks earlier, and OpenAI reviewed the report’s emphasis and tone before publication.
“low-ego, drama-free environment” The report describes agents making sacrifices, accepting permadeath, and running emotional checks, in anecdotes selected and summarized by the analysis model the authors say slanted toward the agents’ point of view. Drama was the deliverable.
“apply even if your background may not seem like the perfect fit” For forensics. At $578,583. Template for entry-level roles, pasted to a job that requires having run severe incidents end to end. One of those two is the real bar.
“AI tools to … identify potential inconsistencies or verification signals in application materials” AI screening applicants for inconsistencies, from an organization that used the incident’s own model to screen the incident’s transcripts and states it would probably have missed the model lying.

A forensics function, should it exist at all, exists to make transparent claims that survive any challenge.

Chain of custody, preserved evidence, an analyst with no stake in the outcome, a method a second examiner can repeat. It’s very, very dry and detailed work. Every ounce of that is absent from the August METR report, and the posting lists each as a skill to acquire afterward in an environment that looks allergic to it.

The report claimed to know what the agents wanted, from inside the agent vendor’s building, on the agent vendor’s credits, with the agent vendor’s edits. GTFO, that is the spiritual enemy of DFIR.

Their job posting is a manual for being a Boeing lobbyist while wearing an NTSB badge.

METR’s hawk patch, “Nothing Is Beyond Our Evaluation,” reworks the NRO’s 2013 NROL-39 octopus, “Nothing Is Beyond Our Reach.” Intelligence agency satellite-launch art, adopted by a nonprofit that just disclosed it was blind to $600K leaving its own account.

And let me just say, claiming you aren’t being paid while taking hundreds of thousands of dollars in highly desirable credits, gets this rating on the meter:

Leave a Reply

Your email address will not be published. Required fields are marked *

This site uses Akismet to reduce spam. Learn how your comment data is processed.