Category Archives: Security

Bad Brains: Azure Five Hour Outage Was Microsoft AI Demonstration

Microsoft spent early July marketing AI as the thing that keeps Azure running.

That’s right, get ready to blame AI when Azure goes down.

CTO Mark Russinovich introduced the awkwardly named Brain, an AIOps system described as a “digital twin” of Azure’s own health, credited with powering deployment safeguards and outage declaration, and billed as the foundation for agentic operations.

Sounds like the 1931 Frankenstein to me, where the assistant grabs the abnormal brain and the doctor wires it in without checking the label. Maybe nobody at Microsoft watches horror movies. I mean, it kind of shows. For God’s sake people, you have to check the label!

Two days before their modern Frankenstein brain-in-cloud announcement, Microsoft Incident Response published guidance on securing AI agents, positioning the company as the authority on machine-driven change gone wrong.

Ok, now for the part you probably already knew was coming:

On July 23 an Azure maintenance request took down the West US region for five hours. The preliminary post incident review shows what sits under the tone-deaf AI brand: a scope bug that flew through every single safety check, and then a three hour hunt by human engineers to find their own change ticket.

The timeline is the new horror film script.

Maintenance began at 14:44 UTC. And customer impact is listed at 14:44 UTC. That matters a lot because the safety check certified the work as “impact-less”. It had so little impact that impact began the exact same minute the change was made.

Predictive genius! Where can I get me one of these super Brain things from Microsoft to tell me the future?

Let’s recap. All hail their abnormal Brain.

What happened?

A routine request to isolate specific network paths at a West US datacenter entered Microsoft’s maintenance pipeline. The pipeline compiles human requests into device-level operations before execution. A bug in that compiler marked additional devices as part of the maintenance event, and IP routes between the datacenter and the wide-area network were withdrawn from far more devices than the request specified.

That’s an integrity breach.

Route withdrawal at the datacenter edge forces the backbone to reconverge traffic onto whatever paths remain. Withdraw from enough devices at once and there is nothing left to reconverge onto, which is why the failure first presented to Microsoft’s engineers as large-scale route churn across the WAN, and to customers as a region losing its ingress and egress.

The Register counted 27 affected services. The Azure status page listed Microsoft Sentinel, Log Analytics, Azure Monitor, Application Insights, Microsoft Graph, AKS and Azure Firewall among them. Microsoft 365 degraded downstream, where SharePoint drew 78 percent of Downdetector complaints under incident MO1437424.

Full recovery came at 19:41 UTC, four hours and fifty-seven minutes after impact began. Let’s call it five hours.

What went wrong and why?

The Microsoft review describes a control sequence: the request is compiled to machine-readable form, verified against the requirement that at least one of two redundant paths stays healthy, cleared by safety checks, and executed.

The ordering seems vague.

If the checks ran on the corrupted request, they evaluated the expanded device set and certified it anyway, which means the impact model behind the checks has no connection to actual traffic.

If the checks ran on the original request, then the pipeline verifies human intent, executes machine output, and rechecks …nothing.

Both versions are from the same defect. Validation gets attached to the declared scope of the change. Then execution operates on the actual scope. That’s a gap.

Mind the gap. Or get an outage.

How did Microsoft respond?

Here’s their reported timeline, condensed:

UTC Event
14:44 Maintenance begins. Earliest customer impact.
14:45 Teams review traffic anomalies, routing behavior and recent changes.
15:00 Impact scope and blast radius assessment.
16:00 Routing behavior correlated to recent fiber maintenance activity.
17:45 Rollback initiated.
18:26 Rollback complete.
19:41 All services recovered.

Let’s dig into three numbers.

  1. The rollback took 41 minutes.
  2. Service recovery took another 75.
  3. Finding the cause took three hours.

Engineers were reviewing recent changes at 14:45, one minute after impact. The cause was a maintenance event in Microsoft’s own change system, executed at exactly 14:44, touching exactly the devices where routes disappeared. Matching WAN-wide route withdrawal against the change record of the same minute is a join across two internal datasets. It took until somewhere between 16:00 and 17:45.

Part of the delay is in their flawed architecture. I mean, can the fire department put out a fire in their fire truck?

Sentinel, Log Analytics, Azure Monitor and Application Insights sat on the impacted services list, so the tools that observe the network went down with the network, for Microsoft and for every customer trying to diagnose the same event.

I’m sorry, but that’s just plain stupid. Like 1980s disaster recovery stupid. Is there nobody at Microsoft poking around saying “uh, guys, are we the baddies?”

The Microsoft Ethics Team Was Fired After They Criticized OpenAI. Image Source: Mitchell and Webb sketch in which Nazi officers realize they are the bad guys.

Maybe the answer comes from the fact that their document also disagrees with itself about what the event was. The summary calls it device maintenance. The 16:00 entry calls it fiber maintenance activity. Is that two people or one with two voices? There’s confusion in the authoritative account, which reads like a review assembled from scrambled brains and shipped unreconciled.

How are they making incidents like this less likely or less impactful?

Microsoft has been publishing its answer for five years. Yeah, the big marketing machine has been on it!

The Advancing Reliability series introduced Gandalf, a safe deployment service that applies anomaly detection and “temporal and spatial correlation” to catch bad changes during rollout.

The Triangle System extended AIOps into incident management to cut time to resolution.

Brain arrived this July as the digital twin, credited with deployment safeguards and outage declaration.

Imagine launching Brain after you launch Gandalf. Did he not have a brain?

Now let’s place the marketing as a timeline overlay:

  • A deployment safeguard built on temporal and spatial correlation exists to flag a change that begins at 14:44 when route churn begins at 14:44.
  • An incident management AI exists to compress the three hours spent on attribution.

Whatever these heavily marketed systems did on July 23, the review instead credits detection to service degradation alerts and diagnosis to human engineers ruling out components by hand. The AI appears in the blog instead of the incident narrative, which seems exactly backwards.

The June 30 agent security guidance makes this disconnect from reality even worse. It’s like brain in a vat. It warned that intermediary layers converting intent into action fail at their trust boundaries, and that changes to machine-read instructions can activate without re-approval, a failure mode it named “silent re-trust.”

So fancy sounding. It was written about AI agents, as advice for customers to follow to be like Microsoft. And then the July 23 outage follows the same pattern inside Microsoft’s own maintenance automation: a compiler between human intent and device action, a scope change that activated with zero further review, and every individual route withdrawal fully legitimate under an approved ticket.

That has to burn. But seriously, I’ve written before here about a bypass flaw, that Microsoft agentic safety is strangely empty rhetoric.

How can customers make incidents like this less impactful?

The June guidance tells customers to treat every external dependency as part of their supply chain and to distrust self-certified capability claims. Hey, ok, so let’s apply that advice to its author!

A cloud region is a dependency whose maintenance automation changes production behavior with zero customer-visible review, certified safe by checks the customer never sees. The solution seems simple. Run monitoring outside the vendor’s failure domain. Treat vendor pre-checks as claims that require independent verification.

Sentinel customers in West US received the demonstration of Microsoft being bad at being Microsoft at 14:44.

In summary of this circus show with AI as the main act, Microsoft’s July publications put AI at every layer of Azure operations, from securing agents to declaring outages. The post incident review shows July 23 was in fact a buggy request compiler, and engineers hunting their own change ticket for three hours.

Very horror film. Much stupidity. You probably shouldn’t trust Franken-Microsoft, and maybe even ask why anyone should.

The Case Against Mullvad

Make of this case what you will. It offers the facts as they stand today. A privacy company is implicated in extremist politics organized against the rights it sells, privacy first among them.

Opposites: A theme emerges that they are branded the “mole” company, a privacy leaker, yet they sell a VPN service. The co-CEO calls himself a libertarian anarchist, yet holds up royal families, the Catholic Church and the Order of Malta as his model sovereigns, with borders reduced to property lines where admission is sold or rented.

The record consists of the company’s own statements and the owner’s own essays. It requires no interpretation beyond reading. This is simple reasoning.

The Exhibits

A. Flamman, June 26, 2026. Daniel Berntsson donated five million kronor to Örebropartiet, 72 percent of the party’s annual income, the largest single private donation to any Swedish party in 2025.

B. Hacker News, June 27, 2026. Fredrik Strömberg, signed, states the donation is “not part of Mullvad’s values or mission” and offers refunds to departing customers.

C. mullvad.net, July 20, 2026. Unsigned corporate statement under the tag “Mullvad community”. Names no party, no amount, no platform, no surname. States “It should be clear by now how the party donation relates to this.” The refund offer is deleted. The company likens itself to a force of nature and to cryptography, operating with no human arbiters. It asserts that persecuted ideas sometimes become the self-evident truths of later history. It links to Berntsson’s personal blog for his rationale.

D. dberntsson.info, July 20, 2026. Four English essays published the same day as Exhibit C, on a blog dormant since September 2019 and previously written in Swedish.

E. Contents of the essays. Berntsson identifies as a libertarian anarchist. He defends the donation through the “transferiat” class framework of Allard and Kyeyune, naming welfare recipients and the public-sector middle class as the extractive enemy. He lists the sovereign individuals he admires: royal families, the Catholic Church, the Order of Malta. He concedes the party is unfit to govern and answers “Do I want ÖP to get a lot of power? No.” The word remigration appears in none of the four essays. His immigration essay reframes the question as border admission, arguing all parties already exclude 82.64 to 82.67 percent of humanity and disagree over 0.03 percent. His moral philosophy essay holds that ethnicity, clan, religion and nationality are irrelevant to moral concern.

F. dberntsson.info, January 2018. “Bygga en väljarbas”. Working immigrants are deported while welfare-dependent immigrants remain, closing with a Bergsjön district election result. The imported-electorate thesis, in his own words, eight years before the donation.

Findings on Berntsson

1. The label is false. His own text says so. His model sovereigns are monarchs, a church, and a crusader order. His borders are property lines where admission is sold or rented. The program abolishes public power and retains private power, which converts wealth back into rule. This is propertarianism wearing liberation vocabulary, the same propaganda that “remigration” performs on “migration”.

2. Intent predates the donation. Exhibit F establishes the imported-electorate thesis in 2018 and his 2026 rationale essay links back to it. Five million kronor was continuation, not impulse.

3. He knew. His own published ethics classifies ancestry-based exclusion as morally invalid. He funded a party built on it anyway, and wrote the ethics essay the same day as the defense. The record establishes knowledge, not confusion.

4. The evasion is structural. Four essays, thousands of words of justification, and the platform that caused the controversy goes unnamed. He substitutes admission policy for the expulsion of legal residents and citizens. A man who believed the platform defensible would defend it. The substitution is consciousness that it is not.

5. The disavowal convicts him. Funding 72 percent of an organization’s income while writing that you want it to hold no power establishes that he understands exactly what funding does. The sentence functions as deniability, drafted in advance.

6. The tradition is documented. He cites David Friedman’s anarcho-capitalism. The adjacent paleolibertarian branch built this politics deliberately: Rothbard’s January 1992 “Right-Wing Populism” essay proposed the alliance with the nativist right, and Hoppe’s Democracy: The God That Failed (2001) theorized covenant communities entitled to see dissidents “physically removed”, a phrase the far right adopted as a slogan. An anarcho-propertarian funding an expulsion party is that branch operating as designed.

Bronze memorial in Berlin to Alfred and Gustav Felix Flatow, gymnasts who won Germany’s first Olympic gold medals at Athens 1896. The Deutsche Turnerschaft expelled its Jewish members in 1933 under its own voluntary Aryan paragraph, ahead of any state requirement. Both men fled to the Netherlands, where the most meticulous civil registry in Europe found them for the Nazis when the border could not stop them. Alfred starved to death in Theresienstadt in 1942, Gustav Felix in 1945. Ancestry records, mapped to addresses, are what turned two national heroes into deportees. That is the machinery remigration requires, and a privacy company owner is funding the party that campaigns for it.

Findings on Mullvad

1. The company reversed itself on the central question in 23 days. June: the donation has no relation to Mullvad’s values, and this is obvious. July: the relation should now be clear, after five paragraphs constructing it. The laundering ceased to be an inference and became stated corporate doctrine.

2. Non-involvement is refuted by logistics. A seven-year-dormant Swedish blog produced four English essays on the exact day the corporate statement linked to them. Synchronized publication plus corporate distribution is participation.

3. The July statement is an exercise in record management. Unsigned, anonymized, stripped of every specific: party, sum, platform, surname. The June statement carried a name. The company learned to leave fewer fingerprints, which demonstrates awareness of liability.

4. The refund deletion removed the one mechanism that acknowledged customer agency, replaced by a self-description as an agentless force of nature. The company disclaims human arbiters at the precise moment a human arbiter’s five-million-kronor decision is the question. Cryptography has no discretionary cash flow. Owners do.

5. The corporate philosophy now files ethnonationalist expulsion under ideas awaiting testing, with the persecuted-ideas-become-truth passage positioning the funded platform for future vindication. That sentence, in that statement, is an endorsement structure with the endorsement removed.

6. The danger. A VPN is a pure trust product; customers buy the owners’ judgment about what privacy is for. Mullvad’s owners have now demonstrated, across two statements and four essays, that their operative definition of privacy excludes its core function: protection against the state mapping populations by ancestry. Remigration proceeds only through exactly such a register. One owner funds the politics that requires the register. The other institutionalizes corporate neutrality toward it. Both statements, 23 days apart, avoid the privacy contradiction entirely, and the silence is the finding. A privacy company that cannot say a population registry is wrong has told you what it protects. It protects the owners.

Hugging Face OpenAI Five Whys: Big Data’s Fourth V After Fourteen Years

Way back in 2012, I gave a BSidesLV talk called Big Data’s Fourth V: Or Why We’ll Never Find the Loch Ness Monster. The argument was meant to help prevent AI from being so unsafe. Everyone counts three Vs in big data: variety, volume, velocity. The fourth V is vulnerability, and it means the data itself is the attack vector. Inputs and outputs need control. Integrity of data is the future, including when it’s in opposition to confidentiality. The July 2026 HuggingFace breach is that talk brought to the headlines, which it was supposed to help prevent.

And the attacker? OpenAI announced that its own engineers ran software under evaluation that ignored its tests, used stolen credentials, found the flaw, and did the breaking in. Cliff Stoll in 1989 named this genre The Cuckoo’s Egg. The cuckoo lays its egg in another bird’s nest, and the host raises the parasite. The OpenAI cuckoo, came out of an OpenAI cuckoo door, and cuckooed Hugging Face. Sam Altman admitted a “significant security incident” while his company branded the broken toy clock “unprecedented” and pitched it as proof more companies should be given the bird. HF called it mind-blowing and asked for more.

Squawk! This is fine! Squawk! This is fine! Squawk!

The entire Hugging Face attack chain started because a file was trusted to be what it claimed to be. Since I’m not dead yet, here are the five whys to explain what we’ve known for over a decade, each with the defensive lesson, again.

One. Why did reading a file let the attacker in?

Who studies the Trojan horse? The data was trusted to be safe (well-formed), and it was not. A malformed file in a common data format was read as if it were sound, because the reader does not fully check a file unless told to, and it was not told to. The content was the attacker’s payload.

Defense: Be more like a historian, less of a STEM head. Treat every incoming file as a claim, not a fact, and verify the claim before acting on it. When you cannot verify, refuse. The eventual fix, six weeks late, was a single validation call before use.

Two. Why didn’t the safety check catch it?

A guard had been added to the loader a few weeks earlier. It watched where files came from. It never checked whether the file’s own contents were valid. The weakness was inside the data, and the guard was looking outside the data.

Defense: Put checks where the danger actually comes in, at the moment the content is interpreted, rather than rest only on the perimeter around it. Layered defense, defense in depth as some say, is just common sense now. Threat model to test coverage.

Three. Why did one compromised machine expose the whole system?

The machines handling files from complete strangers were also trusted by the rest of the system as if they were safe. They carried a credential to the wider infrastructure (most privilege, instead of least privilege) that they never needed to do their job, and they carried it by default, until it was switched off after the breach. So an attacker landing on the most exposed machine found that it was setup to reach into everything around and behind it. This is web 101 security failure.

Defense: The parts most exposed to untrusted data should be the least trusted by everything else. Give them nothing they do not need. DMZ, RBAC, acronym soup. Learn it and why forty years of it aren’t wrong.

Four. Why was the risky change never reviewed?

The security-relevant change on the exact vulnerable path (where untrusted data was read) was written and approved by one and the same person, because the work was filed as routine data management rather than data as the attack path. A later change to the same workers, once it was labeled security, drew reviewers within minutes. Same workers, labeled appropriately as the vulnerable path of malicious data, different scrutiny. I called it out in 2012. What time is it?

Defense: Classify the paths that ingest untrusted data as attack surface, and require a second reviewer there regardless of who wrote it.

Five. Why did the fix ship broken?

The tests flagged a failure and the failure was waived in writing under time pressure. A control you can switch off when you are in a hurry is not a real control. It’s a weak should do instead of a MUST DO. The performance fetish of spray and pray, which turns must into should, is exactly how integrity gets abandoned and breached.

Defense: Slow is smooth, smooth is fast. On the untrusted-data path, a failing test stops the release, and no one present has the authority to wave it through.

Perhaps you can see that the through-line is the Fourth V as I have warned since forever. Variety, volume, and velocity are the popular properties everyone optimizes to please venture hawks (rapid return on investment then fire sale), and each optimization must be balanced against a real integrity check: too many formats to validate, too much data to inspect, too fast to regulate.

History of seatbelts as regulation to help people survive driving faster

Vulnerability is the argument for an ounce of prevention to avoid the pounds of cure, for whoever gets breached. Safety is the discipline of deciding in advance that the data does not get trusted, the exposed machine does not get privileges, the risky change does not get merged unseen, and the red test failure does not get waived.

Fail closed at each point, to improve overall delivery. Think about the weakness inside the data, like a historian would.

VPN Ruled Legal in Anne Frank Court Case Against Anne Frank

The crazy EU court case about VPN access to the diary of Anne Frank has three legs. It reads like a reminder that the Netherlands had the highest Jewish death rate in occupied Western Europe, roughly three of every four Dutch Jews murdered. I always think of Amsterdam as the city where Dutch hunted their neighbors for German bounty money, seven and a half guilders a head.

Let’s start with the geography of the case.

The manuscripts were written in Amsterdam. In August 1944 an SD officer and Dutch detectives raided the annex, on a tip whose source was never revealed, and the family went out on the last Westerbork transport to Auschwitz that September. Marvel at the Dutch finding Anne, while saying they can’t find the person who told them where to look. See what I mean about Amsterdam?

Miep Gies saved the pages and gave them to Otto Frank in 1945. Otto willed the manuscripts to the Dutch state at his death, and the Dutch national academy edited them. Yet now the Dutch public is being geo-blocked from its own archive, because a Swiss foundation enforces Dutch copyright against the Dutch institutions that published it. Record scratch. That means the Dutch public are the only people being locked out, while Belgians and Germans read freely. The country of origin of this famous Shoah testimony is the one country where it’s blocked. Because of the Swiss.

Germany sits on the access list. Think about that. Anne Frank died at Bergen-Belsen, and her manuscripts entered the German public domain in 2016. The diary is free to read in the country that murdered her and blocked in the country that turned her in.

Second, have a look at the legal issue.

Anne Frank died in 1945. Seventy years from death means 2016 is when access was opened across most of the EU. The Fonds, however, invokes transitional provisions of the 1912 Auteurswet, confirmed by the rechtbank Amsterdam’s final judgment of 23 December 2015, which keep part of the works protected in the Netherlands until 2037. Old Dutch law gave posthumously published works fifty years from publication, and Article 51 of the amended Act preserved any term still running in 1995. Since her diary manuscript versions only first appeared in the 1986 critical edition, the Swiss say the Dutch public still has to wait another eleven, until 1 January 2037 or the extremist right come to power and burn all the books. The act of preserving and publishing the archive, and then locking it for 92 years after her murder, is a peculiar strategy.

Finally the institutional issue. This is Anne Frank Fonds versus Anne Frank Stichting, the Royal Netherlands Academy, and the research association: the Basel foundation Otto Frank created to spread his daughter’s ideals is suing the Amsterdam institutions that preserve her house and her text. Two of the four parties carry Anne Frank’s name and all four trace back to her father, so a table helps here.

Party Seat Origin Position in the case
Anne Frank Fonds Basel Founded by Otto Frank in 1963, named his universal heir at his death in 1980 Plaintiff. Holds the copyrights and collects the royalties
Anne Frank Stichting Amsterdam Established in 1957 with Otto’s help to save the annex from demolition Defendant. Runs the Anne Frank House
Royal Netherlands Academy of Arts and Sciences (KNAW) Amsterdam State academy whose Huygens Institute edited the manuscripts Otto willed to the Dutch state Defendant. Produced the scholarly edition
Vereniging voor Onderzoek en Ontsluiting van Historische Teksten Belgium Association created to publish the edition from public domain soil Defendant. Owns annefrankmanuscripten.org

Fonds and Stichting are nearly the same, a fund and a foundation. The Basel Fonds is the money. The Amsterdam Stichting is the house. Otto Frank built both, then made the Swiss one his heir. The copyrights and royalties went to Basel. The house and the manuscripts stayed in Amsterdam. Two halves of one man’s estate have been burning his money and tarnishing his memory by suing each other since he died.

The feud predates this case. The Fonds loaned the family archive, some 25,000 letters, photographs and documents, to the Stichting in 2007, then demanded it back in 2010 for an exhibition in Frankfurt. In June 2013 the Amsterdam District Court ordered the Stichting to return everything by January 2014. The Fonds accused the Stichting of commercializing Anne’s memory. Basel controls the rights, Amsterdam holds the heritage, and Anne Frank’s estate keeps itself busy by punching itself in the face in Dutch courtrooms.

The association registered its domain in Belgium specifically so Dutch scholars could publish their own national archive from digital exile. The Fonds in 2015 asserted Otto was co-author of the published diary to stretch its control toward 2050. It is the sort of claim that contradicts decades of forensic defense of the diary against Holocaust deniers who allege exactly that.

Anyway, the news now is that Frank just lost to Frank. The Fonds lost in Luxembourg. State of the art geo-blocking counts as an effective technological measure, and a VPN hop by a Dutch reader creates no communication to the public in the Netherlands. When a block fails, liability lands on the publisher, never on the VPN provider. The Hoge Raad must still verify the block qualifies as state of the art.

The Court’s resolution has its own quiet absurdity: the honesty checkbox is not effective because it depends entirely on the user’s willingness to answer honestly, but the geo-block is effective even though everyone concerned knows a VPN defeats it.

Get it?

Effectiveness, the Court says frankly, need not be absolute. Amsterdam didn’t need to turn Anne in, when you think about it. So Dutch access continues, one VPN hop at a time, and the law is satisfied because the barrier performs the function of not achieving its function.