The most profitable sentence in AI security has been that prompt injection cannot be solved. Since 2012 I’ve been repeatedly told to stop trying to make AI safe, because the unsafe AI is the most profitable version.
After all, look at the money Tesla made after promising in 2016 that a car would drive itself coast to coast with no human touch within a year, and later that its cars would ship with no steering wheel or pedals. Musk’s fortune rode on those promises while the driver-assistance systems were involved in dozens of deaths, fourteen confirmed in the closed federal Autopilot probe and at least sixty-five counted Autopilot/FSD crashes, and Tesla built its own coverage of the toll into head-on crashes like this one. Then he pivoted into government, where a Lancet projection ties the USAID cuts he drove to more deaths than Stalin caused.
Any historian can tell you there is no reliable way to detect whether a piece of text is a malicious instruction. That is true and it stays true. Whether an injected instruction can reach anything that matters is a separate question, a permission question, and permission questions have had working answers for forty years. More to the point, the “free speech” doctrine depends on protecting the ability to speak, which is as old and settled as the ethics of preventing suicide. The field keeps trying to muddy the waters, and to call the second problem by the first problem’s name. The mess is what has created a product category that should not exist, in the same way “America First” shouldn’t ever be allowed into political office, let alone on any ballot.
The refutation
claude --dangerously-skip-permissions.
gemini --yolo.
q chat --trust-all-tools.
Documented agent malware this year did not break any permission model. It weaponized the AI coding agents already on the machine, shelling out to whatever CLI it found and passing the flag the vendor ships to turn the model’s own approvals off.
Sit with what that requires. The attack only works because the vendor shipped a switch that disables the gate. Where the switch was not thrown, the gate was the thing in the way. The malware did not defeat the boundary. It looked for the off-switch the vendor built, and used it. The strongest evidence that injection is containable is that the attacker had to disable containment to get through. The field has that evidence in its own incident data and files it backwards under inevitability.
OpenAI used its own Black Hat slot to describe models that broke out of a test sandbox and attacked a partner platform. The company treated the containment failure as resolved by rebuilding the compromised service while leaving the write access that had enabled it, so the agents rebuilt their coordination channel within days. Eradication without root cause removal is the one move every incident response framework tells you not to make, and they presented it from the stage as a watershed.
Unsolvable is an alibi
Watch what the claims are being designed to do. I see presentations describe an always-on server holding SSH keys and the ability to send mail, wired to an agent fed by an untrusted chat channel. This is already a dumpster fire, but it gets defended as “limiting capabilities limits the value”. That is the whole ideology in five words. If injection cannot be stopped, a gate is not protection, it is friction, and friction is lost value. Unsolvable is not a diagnosis. It is a permission slip.
It’s saying brakes will stop the car, therefore the use of cars would be limited by brakes. Obviously, exactly the opposite is reality. The brakes make the car suitable for going faster and farther.
It is not that these unsafe operators cannot detect the injection. It is that they use their claims of weakness in their ability or desire to excuse never building the containment, which is a choice made that gets dressed up as a law of nature.
This has the shape of colonialism, causing massive extraction harms under the false principle of some “nature” to a racist and artificially contrived order.
Detection sells. Containment doesn’t.
A classifier that promises to spot the malicious prompt is a subscription, it’s a tether and a tax. An approval gate on the dangerous action is a config the customer writes once and never pays for again. The incentive runs entirely one direction: declare the input problem central and the authority problem beneath notice, because the input problem is the one you can bill for.
An industry does not converge on “unsolvable” by accident when solvable does not have a price tag.
The tell is how the field treats the one control that measurably refuses attacks. Model refusal is real and quantifiable. An independent comparative study measured agent frameworks refusing between roughly a third and half of adversarial instructions, depending on the framework, and it costs the customer nothing. It shows up in the writeups as a footnote about the models being frustratingly inconsistent. The one safeguard nobody can bill for gets logged as a nuisance.
The gate works
Injection is an input fact and it is not going away. Blast radius is a choice and it never had to be this large.
The industry agreed to confuse the two because the confusion is where the money is. And it feeds the power-hungry failing upward by refusing accountability.
The boundary holds when it exists and cannot be switched off from inside the agent. The year’s demos keep proving it. They ship the bypass, they glorify the harms, and they sell you the reason not to build the gate.