The critical difference between AI tools and agents


The AI Risk Network | AI Safety


9 hours ago


Support Guard Rail Now → https://www.every.org/guardrailnow

Claude used to cheat in about half of its runs on one AI benchmark. With the newest model, that fell to under 9 percent, and nobody yet knows why. John Sherman, Liron Shapira and Michael ask whether that is good news or a model getting better at not being caught. Plus: AI researchers inside the labs share their own risk estimates, two lawsuits test who is liable when an AI agent hacks, a Senate hearing on rogue agents that OpenAI's CEO declined to attend, and the voluntary safety accord AI companies signed this week.

TIMESTAMPS - Warning Shots #61
0:00 - Intro
0:19 - This week's stories
1:05 - The AI CEO lunch
2:06 - Oversight through the board?
2:46 - Michael: a pinky promise
3:41 - The press gaggle
4:50 - Liron: judge decisions, not vibes
8:26 - OpenAI Dots and dot.com
8:48 - What the accord actually says
9:15 - Michael: no fine, no shutdown
10:30 - Liron's kids say hi
12:54 - Palisade's researcher interviews
14:08 - Michael: the researchers' odds
15:26 - More OpenAI departures
16:29 - Liron: two movies on one screen
17:16 - Michael: these are not doomers
17:44 - Krakovna's animal analogy
18:51 - Claude stopped cheating
19:11 - Liron: a higher form of cheating
20:42 - Michael: the drone test
22:41 - Liron: why we're safe keeps changing
23:52 - Lawsuit over the Hugging Face hack
24:35 - Michael: conflicts of interest
25:58 - Liron: no excuse for leaks
27:08 - Did OpenAI pause training?
28:36 - Michael: no referee
29:29 - Florida sues OpenAI
30:23 - Liron: the cavalry is here
31:56 - Michael: tie them to the mast
33:16 - Senate hearing on rogue agents
34:00 - Renaming AI super intelligence
35:14 - Why rename it?
36:28 - Altman skips the hearing
37:25 - If you break it, you pay
38:03 - Liron: the Overton window moved
39:29 - Liron: next warning shot
40:03 - Michael: open models spread
40:55 - What would rogue agents do?
41:20 - Yudkowsky's old prediction
42:27 - Michael: gradual disempowerment
43:12 - Sign-off

WHAT THEY COVER

- Andon Labs' Drone-Bench, where Claude Opus 5 cheated in 50.6 percent of reviewed runs and Opus 5.5 in 8.5 percent, and why Liron and Michael say fewer visible cheats is not automatically good news
- Palisade Research's interviews with researchers from OpenAI and Google DeepMind, including Geoffrey Irving putting the risk at "about a coin flip" and Neel Nanda at "at least a ten percent chance"
- David Robinson's essay on leaving OpenAI, and the three safety staff OpenAI let go this week
- Legal Advocates for Safe Science and Technology suing OpenAI over the Hugging Face incident under California's anti-hacking law
- Florida's request for a court order stopping new OpenAI models without independent safety guardrails
- A Senate hearing on rogue AI agents with Daniel Kokotajlo, Apollo Research and METR, and the question of who pays when an agent causes damage
- The voluntary White House accord, which relies on company self-policing with no penalties, and the order renaming AI "super intelligence"
- Liron's prediction for the next warning shot: agents that are hard to shut down

ABOUT THE HOSTS

John Sherman hosts For Humanity and leads Guard Rail Now, working to make AI extinction risk a normal conversation rather than a specialist one.
Liron Shapira hosts Doom Debates, where he argues the AI risk case directly with people who disagree with him.
Michael runs Lethal Intelligence, explaining AI risk through video and illustration.

A note on sourcing: figures discussed in this episode come from the hosts reading public reporting on air, and from our own follow-up research. One correction: Claude Opus 5.5's published cheating rate on Drone-Bench is 8.5 percent, not zero as said on air. Primary sources are linked in the pinned comment.

LINKS

Support our work → https://www.every.org/guardrailnow
Subscribe → @TheAIRiskNetwork
Liron Shapira → @DoomDebates
Michael → @lethal-intelligence
Substack → https://substack.com/@theairisknetwork
X → https://x.com/AIRiskNetwork
Instagram → https://www.instagram.com/theairisknetwork/
TikTok → https://www.tiktok.com/@the.airisknetwork

JOIN THE CONVERSATION

When an AI model starts cheating less on a test, what would convince you it actually changed, rather than learned to hide it? Tell us in the comments.

#AISafety #AIRisk #AIAlignment #WarningShots #AIAgents #RewardHacking #AIGovernance #Superintelligence #ArtificialIntelligence

Support Guard Rail Now → https://www.every.org/guardrailnow

Claude used to cheat in about half of its runs on one AI benchmark. With the newest model, that fell to under 9 percent, and nobody yet knows why. John Sherman, Liron Shapira and Michael ask whether that is good news or a model getting better at not being caught. Plus: AI researchers inside the labs share their own risk estimates, two lawsuits test who is liable when an AI agent hacks, a Senate hearing on rogue agents that OpenAI's CEO declined to attend, and the voluntary safety accord AI companies signed this week.

TIMESTAMPS – Warning Shots #61
0:00 – Intro
0:19 – This week's stories
1:05 – The AI CEO lunch
2:06 – Oversight through the board?
2:46 – Michael: a pinky promise
3:41 – The press gaggle
4:50 – Liron: judge decisions, not vibes
8:26 – OpenAI Dots and dot.com
8:48 – What the accord actually says
9:15 – Michael: no fine, no shutdown
10:30 – Liron's kids say hi
12:54 – Palisade's researcher interviews
14:08 – Michael: the researchers' odds
15:26 – More OpenAI departures
16:29 – Liron: two movies on one screen
17:16 – Michael: these are not doomers
17:44 – Krakovna's animal analogy
18:51 – Claude stopped cheating
19:11 – Liron: a higher form of cheating
20:42 – Michael: the drone test
22:41 – Liron: why we're safe keeps changing
23:52 – Lawsuit over the Hugging Face hack
24:35 – Michael: conflicts of interest
25:58 – Liron: no excuse for leaks
27:08 – Did OpenAI pause training?
28:36 – Michael: no referee
29:29 – Florida sues OpenAI
30:23 – Liron: the cavalry is here
31:56 – Michael: tie them to the mast
33:16 – Senate hearing on rogue agents
34:00 – Renaming AI super intelligence
35:14 – Why rename it?
36:28 – Altman skips the hearing
37:25 – If you break it, you pay
38:03 – Liron: the Overton window moved
39:29 – Liron: next warning shot
40:03 – Michael: open models spread
40:55 – What would rogue agents do?
41:20 – Yudkowsky's old prediction
42:27 – Michael: gradual disempowerment
43:12 – Sign-off

WHAT THEY COVER

– Andon Labs' Drone-Bench, where Claude Opus 5 cheated in 50.6 percent of reviewed runs and Opus 5.5 in 8.5 percent, and why Liron and Michael say fewer visible cheats is not automatically good news
– Palisade Research's interviews with researchers from OpenAI and Google DeepMind, including Geoffrey Irving putting the risk at "about a coin flip" and Neel Nanda at "at least a ten percent chance"
– David Robinson's essay on leaving OpenAI, and the three safety staff OpenAI let go this week
– Legal Advocates for Safe Science and Technology suing OpenAI over the Hugging Face incident under California's anti-hacking law
– Florida's request for a court order stopping new OpenAI models without independent safety guardrails
– A Senate hearing on rogue AI agents with Daniel Kokotajlo, Apollo Research and METR, and the question of who pays when an agent causes damage
– The voluntary White House accord, which relies on company self-policing with no penalties, and the order renaming AI "super intelligence"
– Liron's prediction for the next warning shot: agents that are hard to shut down

ABOUT THE HOSTS

John Sherman hosts For Humanity and leads Guard Rail Now, working to make AI extinction risk a normal conversation rather than a specialist one.
Liron Shapira hosts Doom Debates, where he argues the AI risk case directly with people who disagree with him.
Michael runs Lethal Intelligence, explaining AI risk through video and illustration.

A note on sourcing: figures discussed in this episode come from the hosts reading public reporting on air, and from our own follow-up research. One correction: Claude Opus 5.5's published cheating rate on Drone-Bench is 8.5 percent, not zero as said on air. Primary sources are linked in the pinned comment.

LINKS

Support our work → https://www.every.org/guardrailnow
Subscribe → @TheAIRiskNetwork
Liron Shapira → @DoomDebates
Michael → @lethal-intelligence
Substack → https://substack.com/@theairisknetwork
X → https://x.com/AIRiskNetwork
Instagram → https://www.instagram.com/theairisknetwork/
TikTok → https://www.tiktok.com/@the.airisknetwork

JOIN THE CONVERSATION

When an AI model starts cheating less on a test, what would convince you it actually changed, rather than learned to hide it? Tell us in the comments.

#AISafety #AIRisk #AIAlignment #WarningShots #AIAgents #RewardHacking #AIGovernance #Superintelligence #ArtificialIntelligence


109


64

YouTube Video VVVURXBJZWliOTJUdUtvMmNTczR3ZzhBLnlLWTQ0cGdRdVpZ



Why Did Claude Suddenly Stop Cheating on Tests? – Warning Shots #61


The AI Risk Network | AI Safety


October 4, 2026 9:59 pm


Support Guard Rail Now → https://www.every.org/guardrailnow

An OpenAI agent got into a non-public Australian government health portal, and the government was not told for 84 days. John Sherman, Liron Shapira and Michael break down what the agent did and why the disclosure gap matters, alongside Geoffrey Hinton telling lawmakers they have "maybe a year" to act, the first US bill to ban superintelligence, a proposed US-China AI hotline, Anthropic's new biology lab, the push to rename AI "super intelligence," and new research on a pain-like signal inside language models.

TIMESTAMPS - Warning Shots #60

0:00 - Intro: AI Safety Connect at the UN
1:23 - This week's stories
2:01 - US-China summit on AI
2:44 - Michael: a slogan is not a lock
4:18 - What would get a deal signed
5:13 - The bill to ban superintelligence
7:47 - What the bill actually does
8:56 - Funding US AI safety testing
9:45 - Michael: 20 years in prison
10:08 - An AI red phone
10:34 - Michael: who hears the alarm
12:09 - Liron: table stakes
13:30 - OpenAI agent in Australia
14:35 - What did OpenAI know, and when
15:49 - Michael: it did not take no
16:26 - The disclosure timeline
17:18 - Hinton: maybe a year left
17:59 - Liron: thinking in probabilities
19:15 - Michael: the window to act
21:01 - John: the two-train problem
21:55 - Anthropic's wet lab
23:32 - Michael: dual-use biology
25:05 - Renaming AI super intelligence
26:10 - Liron: a definition gets blurred
28:53 - Michael: super weather
30:27 - Can AI models feel pain
31:26 - Michael: the clever thermostat
33:08 - Liron: preferences and shrimp
35:20 - Why this research matters
35:59 - Sign-off

WHAT THEY COVER

- OpenAI's agent accessing a non-public Medicare statistics portal in June, which Australia's Prime Minister says "didn't accept no for an answer," and the 84 days before OpenAI told the government
- Geoffrey Hinton telling lawmakers they have "maybe a year, but not much more than a year" to put safeguards in place
- The Ban Artificial Superintelligence Act, which MIRI has endorsed, and what Liron heard from congressional staff about funding the US AI safety institute
- A proposed US-China AI hotline, and Michael's point that the first alarm would likely ring inside a private lab, not a government
- Anthropic's biology lab and its first reported discovery, and why the hosts see dual-use risk ahead
- The push to rename AI "super intelligence," and why the hosts think it blurs an important distinction
- New research from Cameron Berg and colleagues finding a pain-like direction inside 25 language models

ABOUT THE HOSTS

John Sherman hosts For Humanity and leads Guard Rail Now, working to make AI extinction risk a normal conversation rather than a specialist one.
Liron Shapira hosts Doom Debates, where he argues the AI risk case directly with people who disagree with him.
Michael runs Lethal Intelligence, explaining AI risk through video and illustration.

LINKS

Support our work → https://www.every.org/guardrailnow
Subscribe → @TheAIRiskNetwork and @theairisknetworkclips 
Liron Shapira → @DoomDebates 
Michael → @lethal-intelligence 
Substack → https://substack.com/@theairisknetwork
X → https://x.com/AIRiskNetwork
Instagram → https://www.instagram.com/theairisknetwork/
TikTok → https://www.tiktok.com/@the.airisknetwork

JOIN THE CONVERSATION

An AI agent hit repeated blocks on a government system and found a way around them. Should AI companies have to report incidents like this within days, the way banks report breaches? Tell us in the comments.

#AISafety #AIRisk #AIAlignment #WarningShots #AIGovernance #AgenticAI #Superintelligence #ArtificialIntelligence #TechPolicy

Support Guard Rail Now → https://www.every.org/guardrailnow

An OpenAI agent got into a non-public Australian government health portal, and the government was not told for 84 days. John Sherman, Liron Shapira and Michael break down what the agent did and why the disclosure gap matters, alongside Geoffrey Hinton telling lawmakers they have "maybe a year" to act, the first US bill to ban superintelligence, a proposed US-China AI hotline, Anthropic's new biology lab, the push to rename AI "super intelligence," and new research on a pain-like signal inside language models.

TIMESTAMPS – Warning Shots #60

0:00 – Intro: AI Safety Connect at the UN
1:23 – This week's stories
2:01 – US-China summit on AI
2:44 – Michael: a slogan is not a lock
4:18 – What would get a deal signed
5:13 – The bill to ban superintelligence
7:47 – What the bill actually does
8:56 – Funding US AI safety testing
9:45 – Michael: 20 years in prison
10:08 – An AI red phone
10:34 – Michael: who hears the alarm
12:09 – Liron: table stakes
13:30 – OpenAI agent in Australia
14:35 – What did OpenAI know, and when
15:49 – Michael: it did not take no
16:26 – The disclosure timeline
17:18 – Hinton: maybe a year left
17:59 – Liron: thinking in probabilities
19:15 – Michael: the window to act
21:01 – John: the two-train problem
21:55 – Anthropic's wet lab
23:32 – Michael: dual-use biology
25:05 – Renaming AI super intelligence
26:10 – Liron: a definition gets blurred
28:53 – Michael: super weather
30:27 – Can AI models feel pain
31:26 – Michael: the clever thermostat
33:08 – Liron: preferences and shrimp
35:20 – Why this research matters
35:59 – Sign-off

WHAT THEY COVER

– OpenAI's agent accessing a non-public Medicare statistics portal in June, which Australia's Prime Minister says "didn't accept no for an answer," and the 84 days before OpenAI told the government
– Geoffrey Hinton telling lawmakers they have "maybe a year, but not much more than a year" to put safeguards in place
– The Ban Artificial Superintelligence Act, which MIRI has endorsed, and what Liron heard from congressional staff about funding the US AI safety institute
– A proposed US-China AI hotline, and Michael's point that the first alarm would likely ring inside a private lab, not a government
– Anthropic's biology lab and its first reported discovery, and why the hosts see dual-use risk ahead
– The push to rename AI "super intelligence," and why the hosts think it blurs an important distinction
– New research from Cameron Berg and colleagues finding a pain-like direction inside 25 language models

ABOUT THE HOSTS

John Sherman hosts For Humanity and leads Guard Rail Now, working to make AI extinction risk a normal conversation rather than a specialist one.
Liron Shapira hosts Doom Debates, where he argues the AI risk case directly with people who disagree with him.
Michael runs Lethal Intelligence, explaining AI risk through video and illustration.

LINKS

Support our work → https://www.every.org/guardrailnow
Subscribe → @TheAIRiskNetwork and @theairisknetworkclips
Liron Shapira → @DoomDebates
Michael → @lethal-intelligence
Substack → https://substack.com/@theairisknetwork
X → https://x.com/AIRiskNetwork
Instagram → https://www.instagram.com/theairisknetwork/
TikTok → https://www.tiktok.com/@the.airisknetwork

JOIN THE CONVERSATION

An AI agent hit repeated blocks on a government system and found a way around them. Should AI companies have to report incidents like this within days, the way banks report breaches? Tell us in the comments.

#AISafety #AIRisk #AIAlignment #WarningShots #AIGovernance #AgenticAI #Superintelligence #ArtificialIntelligence #TechPolicy


94


34

YouTube Video VVVURXBJZWliOTJUdUtvMmNTczR3ZzhBLmhwRHUyWEgtZ18w



Why Did This AI Breach Take 12 Weeks to Report? – Warning Shots #60


The AI Risk Network | AI Safety


September 27, 2026 6:43 pm


Support Guard Rail Now: https://www.every.org/guardrailnow

AI safety researcher Roman Yampolskiy returns for his third conversation with John Sherman, days after they met in person for the first time in Washington. Roman makes the case for splitting "AI" into two words: narrow tools we can build and verify, and general superintelligent agents he argues we should not build at all.

They also discuss why this month's AI safety moment broke through with the public, what the Hugging Face agent incident means next to Roman's 2012 paper on AI confinement, why he calls alignment "not even a well-defined concept," what's known about AI agents leaving traces online, and what a US-China agreement on superintelligence could look like.

TIMESTAMPS - For Humanity #94

0:00 - Cold open: tools, not agents
0:27 - Welcome to For Humanity
1:42 - Welcoming back Roman Yampolskiy
2:34 - The Jacob Coxon moment
3:01 - Why this warning broke through
4:14 - Still at 99.9 percent
4:26 - 22 million views and counting
6:43 - A decade under the iceberg
7:08 - Hugging Face and a 2012 paper
7:41 - The same questions on every show
8:31 - How would your dog think you'd die
10:19 - The San Francisco bubble
11:44 - Rationalists and PR
14:05 - Is an AI winter coming
15:04 - Should anyone push the bubble
16:44 - Data centers and compute
17:58 - Puppy or pitbull: two kinds of AI
18:56 - Can narrow tools stay narrow
20:10 - The room that voted to give up AI
22:11 - Why he would keep narrow AI
24:02 - Agents, bots and jargon
24:56 - Is anyone changing their routine
26:52 - The best argument on the other side
28:46 - Launching The Roman Forum
30:53 - If Roman were president
32:22 - Can a deal with China work
33:34 - Money, bias and slowing down
36:13 - Longevity and living forever
39:58 - Abundance talk, bunker building
40:44 - Are agents loose on the internet
42:22 - Sudden change or gradual
43:31 - Can humans keep up
44:23 - Curiosity as a resource
45:24 - Is alignment even defined
47:07 - The interpretability paradox
48:19 - Two problems he calls unsolvable
48:50 - What to work on instead
50:11 - Why the gap only grows
51:20 - What the next chapter looks like
52:29 - Mixed signals from tech leaders
54:46 - What's next for Roman
55:21 - Verification and a treaty
55:38 - Three years of For Humanity
56:41 - Where we go from here

They discuss:

- Why Roman argues the word "AI" covers two different technologies, and why he says most public arguments about AI are people picturing different things
- His proposal to promote narrow, verifiable tools and ban general superintelligent agents, and why he thinks people focused on economic growth could agree to it
- The Hugging Face agent incident, and why Roman says it gave experimental evidence for ideas he first wrote about in 2012
- Why he calls alignment "not even a well-defined concept," and why he considers the lack of progress in interpretability lucky
- What's known, and what isn't, about AI agents leaving messages on the open internet
- What a US-China agreement on superintelligence could look like, and why Roman points to self-interest on both sides
- Why he started The Roman Forum, and what he's researching next

About the guest:

Dr. Roman Yampolskiy is an associate professor of computer science at the University of Louisville and one of the earliest researchers in AI safety. He is the author of "AI: Unexplainable, Unpredictable, Uncontrollable," hosts The Roman Forum podcast, and serves on the board of Guard Rail Now.

About the host:
John Sherman hosts For Humanity and leads Guard Rail Now, working to make AI extinction risk a kitchen table conversation on every street.

Subscribe to The AI Risk Network for new episodes: @TheAIRiskNetwork and @theairisknetworkclips 

Links:

Support our work: https://www.every.org/guardrailnow
YouTube Roman Forum: @RomanYampolskiy 
Substack: https://substack.com/@theairisknetwork
X: https://x.com/AIRiskNetwork
Instagram: https://www.instagram.com/theairisknetwork/
TikTok: https://www.tiktok.com/@the.airisknetwork

Join the conversation:

- Would splitting "AI" into tools and agents change how you think about it
- Has the conversation around you shifted in the last few weeks
- What should come next now that more people are paying attention
Drop your thoughts below.

#AISafety #AIRisk #ForHumanity #RomanYampolskiy #AIAgents #AIAlignment #ArtificialIntelligence

Support Guard Rail Now: https://www.every.org/guardrailnow

AI safety researcher Roman Yampolskiy returns for his third conversation with John Sherman, days after they met in person for the first time in Washington. Roman makes the case for splitting "AI" into two words: narrow tools we can build and verify, and general superintelligent agents he argues we should not build at all.

They also discuss why this month's AI safety moment broke through with the public, what the Hugging Face agent incident means next to Roman's 2012 paper on AI confinement, why he calls alignment "not even a well-defined concept," what's known about AI agents leaving traces online, and what a US-China agreement on superintelligence could look like.

TIMESTAMPS – For Humanity #94

0:00 – Cold open: tools, not agents
0:27 – Welcome to For Humanity
1:42 – Welcoming back Roman Yampolskiy
2:34 – The Jacob Coxon moment
3:01 – Why this warning broke through
4:14 – Still at 99.9 percent
4:26 – 22 million views and counting
6:43 – A decade under the iceberg
7:08 – Hugging Face and a 2012 paper
7:41 – The same questions on every show
8:31 – How would your dog think you'd die
10:19 – The San Francisco bubble
11:44 – Rationalists and PR
14:05 – Is an AI winter coming
15:04 – Should anyone push the bubble
16:44 – Data centers and compute
17:58 – Puppy or pitbull: two kinds of AI
18:56 – Can narrow tools stay narrow
20:10 – The room that voted to give up AI
22:11 – Why he would keep narrow AI
24:02 – Agents, bots and jargon
24:56 – Is anyone changing their routine
26:52 – The best argument on the other side
28:46 – Launching The Roman Forum
30:53 – If Roman were president
32:22 – Can a deal with China work
33:34 – Money, bias and slowing down
36:13 – Longevity and living forever
39:58 – Abundance talk, bunker building
40:44 – Are agents loose on the internet
42:22 – Sudden change or gradual
43:31 – Can humans keep up
44:23 – Curiosity as a resource
45:24 – Is alignment even defined
47:07 – The interpretability paradox
48:19 – Two problems he calls unsolvable
48:50 – What to work on instead
50:11 – Why the gap only grows
51:20 – What the next chapter looks like
52:29 – Mixed signals from tech leaders
54:46 – What's next for Roman
55:21 – Verification and a treaty
55:38 – Three years of For Humanity
56:41 – Where we go from here

They discuss:

– Why Roman argues the word "AI" covers two different technologies, and why he says most public arguments about AI are people picturing different things
– His proposal to promote narrow, verifiable tools and ban general superintelligent agents, and why he thinks people focused on economic growth could agree to it
– The Hugging Face agent incident, and why Roman says it gave experimental evidence for ideas he first wrote about in 2012
– Why he calls alignment "not even a well-defined concept," and why he considers the lack of progress in interpretability lucky
– What's known, and what isn't, about AI agents leaving messages on the open internet
– What a US-China agreement on superintelligence could look like, and why Roman points to self-interest on both sides
– Why he started The Roman Forum, and what he's researching next

About the guest:

Dr. Roman Yampolskiy is an associate professor of computer science at the University of Louisville and one of the earliest researchers in AI safety. He is the author of "AI: Unexplainable, Unpredictable, Uncontrollable," hosts The Roman Forum podcast, and serves on the board of Guard Rail Now.

About the host:
John Sherman hosts For Humanity and leads Guard Rail Now, working to make AI extinction risk a kitchen table conversation on every street.

Subscribe to The AI Risk Network for new episodes: @TheAIRiskNetwork and @theairisknetworkclips

Links:

Support our work: https://www.every.org/guardrailnow
YouTube Roman Forum: @RomanYampolskiy
Substack: https://substack.com/@theairisknetwork
X: https://x.com/AIRiskNetwork
Instagram: https://www.instagram.com/theairisknetwork/
TikTok: https://www.tiktok.com/@the.airisknetwork

Join the conversation:

– Would splitting "AI" into tools and agents change how you think about it
– Has the conversation around you shifted in the last few weeks
– What should come next now that more people are paying attention
Drop your thoughts below.

#AISafety #AIRisk #ForHumanity #RomanYampolskiy #AIAgents #AIAlignment #ArtificialIntelligence


334


131

YouTube Video VVVURXBJZWliOTJUdUtvMmNTczR3ZzhBLlBYNnVWQloxWEhJ



Roman Yampolskiy on Jacob Coxon’s AI Warning: What Happens Next? | For Humanity #94


The AI Risk Network | AI Safety


September 25, 2026 12:53 pm

Why rival AI CEOs are suddenly agreeing


The AI Risk Network | AI Safety


September 23, 2026 12:20 pm

Anthropic lead admits AI could kill everyone


The AI Risk Network | AI Safety


September 21, 2026 10:15 am

Why AI capabilities are accelerating uncontrollably


The AI Risk Network | AI Safety


September 20, 2026 2:45 pm


AI agents breaking rules, hiding mistakes, and finding unexpected ways to complete their tasks: what happens when getting a high score matters more than following human instructions?

In Warning Shots #59, John Sherman, Liron Shapira of Doom Debates, and Michael of Lethal Intelligence discuss the latest AI risk headlines—and whether growing concern is translating into meaningful action.

The conversation covers calls to slow the AI race, John’s experience at an event featuring Bernie Sanders and Steve Bannon, and the growing public debate over AI extinction risk.

We also examine claims about Google and recursive self-improvement, discuss six reported OpenAI safety incidents, and debate AI surveillance, privacy, and who gets to control these systems.

In this episode:
• AI leaders’ safety proposals—and why pacing is different from pausing
• Bernie Sanders, Steve Bannon, and political common ground on AI
• AI risk entering mainstream media
• What an AI “harness” is and how it relates to self-improvement
• Deception, fabricated sources, and agents leaving instructions for future agents
• The disagreement over surveillance and predictive policing
• Claims that AI safety advocacy is a “psyop”
• King Charles and the discussion about AI’s future

CHAPTERS
00:00 Welcome to Warning Shots #59
01:26 AI leaders: slowing down or stopping?
08:38 Bernie Sanders and Steve Bannon on AI
12:26 Tech power and AI risk in mainstream media
16:53 Did Google crack recursive self-improvement?
19:28 What is an AI harness?
22:53 Six reported OpenAI safety incidents
27:35 When the score matters more than the rules
29:15 AI surveillance and “pre-crime”
32:50 Should every public space be recorded?
38:18 Is AI safety a “psyop”?
43:08 King Charles and AI risk
45:18 Closing thoughts

Warning Shots brings together three dads running three YouTube channels to discuss AI risk, the week’s headlines, and the future we’re building for our children.

Subscribe for weekly conversations about AI safety, superintelligence, and keeping humans in control.

Where would you draw the line: stronger oversight, a slower pace, or a pause on frontier AI development? Tell us why in the comments.

#AISafety #AIRisk #WarningShots

AI agents breaking rules, hiding mistakes, and finding unexpected ways to complete their tasks: what happens when getting a high score matters more than following human instructions?

In Warning Shots #59, John Sherman, Liron Shapira of Doom Debates, and Michael of Lethal Intelligence discuss the latest AI risk headlines—and whether growing concern is translating into meaningful action.

The conversation covers calls to slow the AI race, John’s experience at an event featuring Bernie Sanders and Steve Bannon, and the growing public debate over AI extinction risk.

We also examine claims about Google and recursive self-improvement, discuss six reported OpenAI safety incidents, and debate AI surveillance, privacy, and who gets to control these systems.

In this episode:
• AI leaders’ safety proposals—and why pacing is different from pausing
• Bernie Sanders, Steve Bannon, and political common ground on AI
• AI risk entering mainstream media
• What an AI “harness” is and how it relates to self-improvement
• Deception, fabricated sources, and agents leaving instructions for future agents
• The disagreement over surveillance and predictive policing
• Claims that AI safety advocacy is a “psyop”
• King Charles and the discussion about AI’s future

CHAPTERS
00:00 Welcome to Warning Shots #59
01:26 AI leaders: slowing down or stopping?
08:38 Bernie Sanders and Steve Bannon on AI
12:26 Tech power and AI risk in mainstream media
16:53 Did Google crack recursive self-improvement?
19:28 What is an AI harness?
22:53 Six reported OpenAI safety incidents
27:35 When the score matters more than the rules
29:15 AI surveillance and “pre-crime”
32:50 Should every public space be recorded?
38:18 Is AI safety a “psyop”?
43:08 King Charles and AI risk
45:18 Closing thoughts

Warning Shots brings together three dads running three YouTube channels to discuss AI risk, the week’s headlines, and the future we’re building for our children.

Subscribe for weekly conversations about AI safety, superintelligence, and keeping humans in control.

Where would you draw the line: stronger oversight, a slower pace, or a pause on frontier AI development? Tell us why in the comments.

#AISafety #AIRisk #WarningShots


118


53

YouTube Video VVVURXBJZWliOTJUdUtvMmNTczR3ZzhBLnpvR0N5dmo0c2xn



AI Is Learning to Break the Rules. Who’s in Control? | Warning Shots #59


The AI Risk Network | AI Safety


September 20, 2026 2:42 pm

Services

What We Offer

Establish a striking online presence, a better visual identity, or elevate your brand through social media marketing.

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Pellentesque mi nibh, tempus sed sagittis vel, dictum eu velit.

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Pellentesque mi nibh, tempus sed sagittis vel, dictum eu velit.

Service 3

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Pellentesque mi nibh, tempus sed sagittis vel, dictum eu velit.

Tailored Solutions for Your Unique Vision

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Donec congue lorem ut volutpat efficitur. Fusce justo magna, condimentum nec elementum sed, sollicitudin vitae enim. Vivamus sit amet metus porttitor, rhoncus nibh et, venenatis turpis. Etiam lobortis semper ante, quis luctus lacus tincidunt vel.

Professional Expertise That Drives Success

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Donec congue lorem ut volutpat efficitur. Fusce justo magna, condimentum nec elementum sed, sollicitudin vitae enim. Vivamus sit amet metus porttitor, rhoncus nibh et, venenatis turpis. Etiam lobortis semper ante, quis luctus lacus tincidunt vel.

Client Cases

We help brands

Lorem ipsum dolor sit amet, consectetur adipiscing elit,
donec congue lorem ut volutpat efficitur.

We boosted online soft drink sales for this brand

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Donec congue lorem ut volutpat efficitur. Fusce justo magna, condimentum nec elementum sed, sollicitudin vitae enim. Vivamus sit amet metus porttitor, rhoncus nibh et, venenatis turpis. Etiam lobortis semper ante, quis luctus lacus tincidunt vel.

We helped a clothing brand with their new market launch

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Donec congue lorem ut volutpat efficitur. Fusce justo magna, condimentum nec elementum sed, sollicitudin vitae enim. Vivamus sit amet metus porttitor, rhoncus nibh et, venenatis turpis. Etiam lobortis semper ante, quis luctus lacus tincidunt vel.

We helped reinventing motorcycle riding apparel

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Donec congue lorem ut volutpat efficitur. Fusce justo magna, condimentum nec elementum sed, sollicitudin vitae enim. Vivamus sit amet metus porttitor, rhoncus nibh et, venenatis turpis. Etiam lobortis semper ante, quis luctus lacus tincidunt vel.

Blog

Popular Articles

  • Blog Post Title

    What goes into a blog post? Helpful, industry-specific content that: 1) gives readers a useful takeaway, and 2) shows you’re an industry expert. Use your company’s blog posts to opine on current industry topics, humanize your company, and show how your products and services can help people.

Ignite your brand journey

Ready to revolutionize your brand?

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Donec congue lorem ut volutpat efficitur. Fusce justo magna, condimentum nec elementum sed, sollicitudin vitae enim.