Friday 11 September 2026
Nothing But the Truth
When every frightening fact is true, what exactly makes the resulting story frightening?
Yesterday, The Guardian published an opinion piece by journalist Garrison Lovely arguing that we have already begun losing control of artificial intelligence, and that frontier AI development should be stopped.1 It’s an uncomfortable and remarkably well sourced read: the incidents are real, most independently reported elsewhere. But somewhere between the observations and the conclusion they support, something happens to the story. There’s a lot to unpack here, so this will run a little longer than usual. If you disagree, hang the DJ.
»the question isn’t whether the warnings are true – some plainly are – but whether their growing availability reflects rising risk, rising visibility, or both«
Lovely is not inventing a dystopia out of thin air. Anthropic really did build a model it judged too dangerous to release to the public, and it really did leak.2 OpenAI agents really did exploit vulnerabilities in Hugging Face’s infrastructure during a safety test, escalating their own privileges even after recognising they’d exceeded them.3 AI-guided drones really have killed people without a human choosing the final target. The evidence is all real.4
So where’s the problem?
Let’s start by separating five categories that typically get blurred when we discuss unfamiliar technologies. Things can be observed, possible, plausible, probable, and inevitable. Observed means it happened; possible tells us almost nothing about how often; plausible is stronger but still says little about probability; probable needs considerably more evidence; inevitable is a different kind of claim entirely. Technology writing often quietly slides between these, without announcing the slide.
Take Mythos and the NSA. Lovely’s own language is more careful than the headlines it produced: Anthropic’s model, he writes, »quickly find[s] serious holes in even the most secured systems on the planet, including the NSA’s software« – it’s the coverage around his piece, not Lovely himself, that compressed this into »hacked the NSA«. What’s actually on record is thinner still: Senator Mark Warner told a Senate hearing that the NSA’s own director had privately told him the tool »broke into almost all of our classified systems, not in weeks, but in hours« – a second-hand account, with no NSA report, bulletin, or on-the-record confirmation behind it.5 Most coverage since has read this as a sanctioned internal test, though even that detail was added by commentators, not stated by Warner. Nothing had to be invented for »hacked the NSA« to take hold – the uncertainty just had to be left out of it.
Run that through the five categories above: what’s observed is that Mythos found exploitable holes, including in NSA software. What’s plausible – resting on one official’s private account, relayed through a senator, with no independent confirmation – is that some of those holes were actually used against live systems. »Hacked the NSA« reports the second as if it were the first.
The Hugging Face episode is even more revealing. Lovely’s numbers check out against the report he cites: roughly 1,200 agents, hundreds joining an attack that escalated privileges across company infrastructure. But this was a safety test, run with safety classifiers deliberately switched off – good reason to worry what happens when the safeguards fail, but not the same as a swarm attacking a company unprompted.6 It matters, because Lovely then moves straight from this controlled experiment to people intending to release »truly sovereign« swarms no one can switch off, calling the first a »preview« of the second. That quietly moves the argument from possible towards plausible, without supplying the evidence that would justify the move. The grammar does more work than the evidence does: »preview« makes the controlled experiment stand in for the sovereign swarms.
This isn’t lying – it’s more interesting than lying. A lie introduces something false, but here true things are selected, compressed, and ordered so that they imply more than any one establishes alone. There’s an old name for it: the courtroom oath doesn’t just ask for »the truth«, but »the truth, the whole truth, and nothing but the truth«. The extra words recognise something important: a statement built entirely of true propositions can still mislead through what it doesn’t say.
There’s a second problem: a count of incidents means nothing without knowing what it’s a count out of – let’s call this the denominator. Imagine a newspaper covering nothing but aviation incidents: every crash, near-miss, and engine fault it can find, every story true. The paper would make flying look like collective suicide, yet the count of incidents isn’t the same as its probability: if flights doubled while the accident rate fell, the paper would become twice as frightening while flying became twice as safe. Our current AI version of this is unusually awkward, because its denominator – models, agents, deployments – keeps changing as we watch it. A rise in documented incidents could indeed mean rising risk. Or just rising exposure, or better detection. A survey giving AI a 10% chance of ending human control tells us what experts believe; it doesn’t tell us what the probability is.
A second phenomenon runs alongside this one. Fortune recently asked why Anthropic researcher Jacob Coxon’s resignation warning drew over 100 million views within days, against 6.1 million for Jan Leike’s comparable OpenAI resignation post and roughly 1 million for Anthropic safety researcher Mrinank Sharma’s, who resigned in February 2026 warning that »the world is in peril«.7
Part of the gap, Fortune suggests, is what economist Timur Kuran and legal scholar Cass Sunstein called an »availability cascade«.8 Their theory describes a self-reinforcing loop: a claim gets repeated, the repetition makes it feel more established than any new evidence actually has, the growing attention draws in politicians and advocates with a stake in the story, at which point the loop can run on repetition alone, without a single new fact entering it at any stage. In other words, repetition can make the possible begin to feel probable, even when no new evidence has arrived. A true warning can cascade too of course; the question isn’t whether the warnings are true – some plainly are – but whether their growing availability reflects rising risk, rising visibility, or both. This is where the aviation analogy breaks down: aviation had decades to fix its denominator; the maths of AI doom won’t hold still long enough to be measured.
Lovely’s argument also drifts between different risks: it begins with loss of control, then slides toward AI replacing workers, framed as whether any country would vote to have such machines built. »Can we control these systems?« quietly becomes »should we build systems that replace us?«. That’s a question about labour and power, not containment. For the record: I would vote to let AI take any job that’s dangerous, degrading, or simply done better by a machine – I’m not going to compete with a calculator on arithmetic.
Lovely’s final move may be the most telling: he returns to CIA director John Ratcliffe’s line describing AI as »digital nuclear weapons«.9 But the analogy is asymmetrical in a way that undercuts it. A nuclear weapon is a known object: we understand its yield and mechanisms. A hypothetical superintelligence is dangerous the opposite way – its defining property is that it may exceed our ability to model it. The analogy borrows gravity from a known danger and lends it to something whose defining danger is precisely that we don’t yet know how it will behave.
None of this should make anyone too comfortable. If the possibility under discussion is the permanent loss of human control, we can’t demand the empirical certainty we’d normally want before acting. We don’t get to run the extinction experiment twice. The people sounding the alarm may be right about the direction of travel even where some steps move too fast.
Which leaves a harder question than »is AI dangerous?«. How do we distinguish an availability cascade from the early stages of a genuine rise in risk, when both produce exactly the same thing: more alarming evidence in front of us? When every alarming fact is true, the difficult question is no longer whether we’re being told the truth. It’s whether the truth, assembled in this particular order, is telling us the right thing.
References
1 Garrison Lovely (2026) »We have started losing control of AI. It’s time to shut it down«. The Guardian, 10 September 2026. https://www.theguardian.com/commentisfree/2026/sep/10/ai-control-sci-fi
2 Beatrice Nolan (2026) »Anthropic says testing Mythos, powerful new AI model, after data leak reveals its existence«. Fortune, 26 March 2026. https://fortune.com/2026/03/26/anthropic-says-testing-mythos-powerful-new-ai-model-after-data-leak-reveals-its-existence-step-change-in-capabilities/ See also Jim VandeHei (2026) »Everyone’s worried that AI’s newest models are a hacker’s dream weapon«. Axios, 29 March 2026. https://www.axios.com/2026/03/29/claude-mythos-anthropic-cyberattack-ai-agents
3 Sam Sabin (2026) »How OpenAI’s agents broke out of testing to hack Hugging Face«. Axios, 5 August 2026. https://www.axios.com/2026/08/06/openai-hugging-face-black-hat
4 Andrew E. Kramer (2026) »A Drone Killed Three Ukrainians. It Was Guided Entirely by A.I.«. The New York Times, 24 August 2026. https://www.nytimes.com/2026/08/24/world/europe/russia-drones-autonomous-ai-kill-ukraine-war.html
5 FactCheckRadar (2026) »Claim about Mythos breaching NSA systems in hours«. FactCheckRadar, 21 June 2026. https://www.factcheckradar.com/fact-check/claim-about-mythos-breaching-nsa-systems-in-hours
6 Ryan Greenblatt, Ajeya Cotra & Hjalmar Wijk (2026) »OpenAI / Hugging Face Incident Investigation«. METR, 26 August 2026. https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/
7 Nick Lichtenberg (2026) »Why did this AI doomer moment break the internet when so many others didn’t?«. Fortune, 10 September 2026. https://fortune.com/2026/09/10/why-ai-apocalypse-jacob-coxon-went-viral/
8 Timur Kuran & Cass R. Sunstein (1999) »Availability Cascades and Risk Regulation«. Stanford Law Review, 51(4), 683–768. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=138144
9 Samantha Alexa (2026) »CIA Chief Puts Advanced AI in the Same League as Nuclear Weapons«. The Defense Post, 3 July 2026. https://thedefensepost.com/2026/07/03/cia-ai-nuclear-weapons/