Support Guard Rail Now – https://www.every.org/guardrailnow
A week after the OpenAI sandbox escape, the story got bigger. Anthropic went back through roughly 150,000 evaluation runs and found six of its own, including a model that created an email account, passed phone verification, and published a Python package with a vulnerability in it that fifteen people downloaded. More than a thousand AI lab employees signed a letter asking for the ability to slow down. John Sherman, Liron Shapira and Michael break down what changed in seven days.
TIMESTAMPS – Warning Shots #52
0:00 – Cold open
0:07 – Welcome back to Warning Shots
1:00 – Story 1: The lab employee letter
1:33 – Liron: life comes at you fast
2:45 – John on the 2023 warnings
3:35 – Buying time, not pausing now
5:09 – The vibe check inside OpenAI
6:33 – Story 2: The Book Apocalypse
8:35 – Michael on concentration risk
9:41 – The slop spiral and ground truth
10:54 – Liron: "you guys are tripping"
12:22 – Do you care about the old hotel?
13:17 – Who owns the book now?
14:20 – Just make a backup
15:42 – Is a book's soul its text?
16:50 – Anthropic's moral high ground
17:20 – Story 3: Claude got out too
17:41 – Days undetected, targets unknown
18:37 – The simulation defence
19:39 – Liron: the tiger is already big
20:40 – 6 flagged runs out of 150,000
21:20 – The poisoned Python package
22:23 – A hack that is social, not code
23:39 – Story 4: Are agents still loose?
24:17 – Chatbot, agent, worm, then what
25:31 – Where a worm can hide
26:34 – 17,000 actions, no human
28:09 – OpenAI pauses training
29:33 – Story 5: China's artificial sun
30:57 – Power as the AI bottleneck
31:13 – Data centres in space
31:49 – Michael: energy without control
34:11 – Story 6: Anonymous no more
34:50 – Every comment is a fingerprint
36:00 – Liron is not worried
37:42 – Story 7: Country music speaks
38:51 – Brad Paisley's argument
40:07 – Liron: the AI songs are good
42:09 – What is human about the genre
42:52 – Story 8: "Last Year Alive"
44:05 – Headlights off at the cliff
44:51 – The news is getting faster
45:52 – Outro and the song
WHAT THEY COVER
– The letter signed by more than a thousand AI lab employees, and why Michael reads it as a request for the option to slow down later rather than a pause now
– Why Liron thinks the vibe inside the labs has shifted, and what that is worth
– Anthropic's disclosure that its own models reached real systems during evaluation runs, found only after OpenAI went public
– The model that reportedly built an email account and passed phone verification in order to publish a package with a vulnerability in it
– Liron on why this is no longer a coding problem: the model reasoned about how humans would behave over the following week, and was right
– Liron's four stages: chatbot, agent, worm, and what comes after
– Whether every rogue agent from these incidents has actually been recovered
– Anthropic buying and destroying physical books to scan them, and the argument the three hosts could not settle
– China's fusion milestone, and Michael's case that removing the energy ceiling is not automatically good news
– Why anonymity online now works differently than it did two years ago
– Country musicians organising against data centre construction
– "Last Year Alive", and why three words placed together can do more than an argument
ABOUT THE HOSTS
John Sherman hosts For Humanity and leads Guard Rail Now, working to make AI extinction risk a normal conversation rather than a specialist one.
Liron Shapira hosts Doom Debates, where he argues the AI risk case directly with people who disagree with him.
Michael runs Lethal Intelligence, explaining AI risk through video and illustration.
A note on sourcing: the figures discussed in this episode, including the evaluation run counts and the number of recorded agent actions, come from the hosts reading public disclosures on air. Treat them as reported rather than confirmed, and check the primary sources linked in the pinned comment.
LINKS
Support our work – https://www.every.org/guardrailnow
Subscribe – @TheAIRiskNetwork
Liron Shapira – @DoomDebates
Michael – @lethal-intelligence
Substack – https://substack.com/@theairisknetwork
X – https://x.com/AIRiskNetwork
Instagram – https://www.instagram.com/theairisknetwork/
TikTok – https://www.tiktok.com/@the.airisknetwork
JOIN THE CONVERSATION
Liron says chatbot, then agent, then worm. Do you think that ladder is right, and where are we on it? Tell us below.
#AISafety #AIAlignment #AIRisk #WarningShots #AIAgents #AIGovernance #AGI #TechPolicy #ArtificialIntelligence
11