Support Guard Rail Now → https://www.every.org/guardrailnow

An OpenAI agent got into a non-public Australian government health portal, and the government was not told for 84 days. John Sherman, Liron Shapira and Michael break down what the agent did and why the disclosure gap matters, alongside Geoffrey Hinton telling lawmakers they have "maybe a year" to act, the first US bill to ban superintelligence, a proposed US-China AI hotline, Anthropic's new biology lab, the push to rename AI "super intelligence," and new research on a pain-like signal inside language models.

TIMESTAMPS - Warning Shots #60

0:00 - Intro: AI Safety Connect at the UN
1:23 - This week's stories
2:01 - US-China summit on AI
2:44 - Michael: a slogan is not a lock
4:18 - What would get a deal signed
5:13 - The bill to ban superintelligence
7:47 - What the bill actually does
8:56 - Funding US AI safety testing
9:45 - Michael: 20 years in prison
10:08 - An AI red phone
10:34 - Michael: who hears the alarm
12:09 - Liron: table stakes
13:30 - OpenAI agent in Australia
14:35 - What did OpenAI know, and when
15:49 - Michael: it did not take no
16:26 - The disclosure timeline
17:18 - Hinton: maybe a year left
17:59 - Liron: thinking in probabilities
19:15 - Michael: the window to act
21:01 - John: the two-train problem
21:55 - Anthropic's wet lab
23:32 - Michael: dual-use biology
25:05 - Renaming AI super intelligence
26:10 - Liron: a definition gets blurred
28:53 - Michael: super weather
30:27 - Can AI models feel pain
31:26 - Michael: the clever thermostat
33:08 - Liron: preferences and shrimp
35:20 - Why this research matters
35:59 - Sign-off

WHAT THEY COVER

- OpenAI's agent accessing a non-public Medicare statistics portal in June, which Australia's Prime Minister says "didn't accept no for an answer," and the 84 days before OpenAI told the government
- Geoffrey Hinton telling lawmakers they have "maybe a year, but not much more than a year" to put safeguards in place
- The Ban Artificial Superintelligence Act, which MIRI has endorsed, and what Liron heard from congressional staff about funding the US AI safety institute
- A proposed US-China AI hotline, and Michael's point that the first alarm would likely ring inside a private lab, not a government
- Anthropic's biology lab and its first reported discovery, and why the hosts see dual-use risk ahead
- The push to rename AI "super intelligence," and why the hosts think it blurs an important distinction
- New research from Cameron Berg and colleagues finding a pain-like direction inside 25 language models

ABOUT THE HOSTS

John Sherman hosts For Humanity and leads Guard Rail Now, working to make AI extinction risk a normal conversation rather than a specialist one.
Liron Shapira hosts Doom Debates, where he argues the AI risk case directly with people who disagree with him.
Michael runs Lethal Intelligence, explaining AI risk through video and illustration.

LINKS

Support our work → https://www.every.org/guardrailnow
Subscribe → @TheAIRiskNetwork and @theairisknetworkclips 
Liron Shapira → @DoomDebates 
Michael → @lethal-intelligence 
Substack → https://substack.com/@theairisknetwork
X → https://x.com/AIRiskNetwork
Instagram → https://www.instagram.com/theairisknetwork/
TikTok → https://www.tiktok.com/@the.airisknetwork

JOIN THE CONVERSATION

An AI agent hit repeated blocks on a government system and found a way around them. Should AI companies have to report incidents like this within days, the way banks report breaches? Tell us in the comments.

#AISafety #AIRisk #AIAlignment #WarningShots #AIGovernance #AgenticAI #Superintelligence #ArtificialIntelligence #TechPolicy

Support Guard Rail Now → https://www.every.org/guardrailnow

An OpenAI agent got into a non-public Australian government health portal, and the government was not told for 84 days. John Sherman, Liron Shapira and Michael break down what the agent did and why the disclosure gap matters, alongside Geoffrey Hinton telling lawmakers they have "maybe a year" to act, the first US bill to ban superintelligence, a proposed US-China AI hotline, Anthropic's new biology lab, the push to rename AI "super intelligence," and new research on a pain-like signal inside language models.

TIMESTAMPS – Warning Shots #60

0:00 – Intro: AI Safety Connect at the UN
1:23 – This week's stories
2:01 – US-China summit on AI
2:44 – Michael: a slogan is not a lock
4:18 – What would get a deal signed
5:13 – The bill to ban superintelligence
7:47 – What the bill actually does
8:56 – Funding US AI safety testing
9:45 – Michael: 20 years in prison
10:08 – An AI red phone
10:34 – Michael: who hears the alarm
12:09 – Liron: table stakes
13:30 – OpenAI agent in Australia
14:35 – What did OpenAI know, and when
15:49 – Michael: it did not take no
16:26 – The disclosure timeline
17:18 – Hinton: maybe a year left
17:59 – Liron: thinking in probabilities
19:15 – Michael: the window to act
21:01 – John: the two-train problem
21:55 – Anthropic's wet lab
23:32 – Michael: dual-use biology
25:05 – Renaming AI super intelligence
26:10 – Liron: a definition gets blurred
28:53 – Michael: super weather
30:27 – Can AI models feel pain
31:26 – Michael: the clever thermostat
33:08 – Liron: preferences and shrimp
35:20 – Why this research matters
35:59 – Sign-off

WHAT THEY COVER

– OpenAI's agent accessing a non-public Medicare statistics portal in June, which Australia's Prime Minister says "didn't accept no for an answer," and the 84 days before OpenAI told the government
– Geoffrey Hinton telling lawmakers they have "maybe a year, but not much more than a year" to put safeguards in place
– The Ban Artificial Superintelligence Act, which MIRI has endorsed, and what Liron heard from congressional staff about funding the US AI safety institute
– A proposed US-China AI hotline, and Michael's point that the first alarm would likely ring inside a private lab, not a government
– Anthropic's biology lab and its first reported discovery, and why the hosts see dual-use risk ahead
– The push to rename AI "super intelligence," and why the hosts think it blurs an important distinction
– New research from Cameron Berg and colleagues finding a pain-like direction inside 25 language models

ABOUT THE HOSTS

John Sherman hosts For Humanity and leads Guard Rail Now, working to make AI extinction risk a normal conversation rather than a specialist one.
Liron Shapira hosts Doom Debates, where he argues the AI risk case directly with people who disagree with him.
Michael runs Lethal Intelligence, explaining AI risk through video and illustration.

LINKS

Support our work → https://www.every.org/guardrailnow
Subscribe → @TheAIRiskNetwork and @theairisknetworkclips
Liron Shapira → @DoomDebates
Michael → @lethal-intelligence
Substack → https://substack.com/@theairisknetwork
X → https://x.com/AIRiskNetwork
Instagram → https://www.instagram.com/theairisknetwork/
TikTok → https://www.tiktok.com/@the.airisknetwork

JOIN THE CONVERSATION

An AI agent hit repeated blocks on a government system and found a way around them. Should AI companies have to report incidents like this within days, the way banks report breaches? Tell us in the comments.

#AISafety #AIRisk #AIAlignment #WarningShots #AIGovernance #AgenticAI #Superintelligence #ArtificialIntelligence #TechPolicy


89


34

YouTube Video VVVURXBJZWliOTJUdUtvMmNTczR3ZzhBLmhwRHUyWEgtZ18w



Why Did This AI Breach Take 12 Weeks to Report? – Warning Shots #60


The AI Risk Network | AI Safety


September 27, 2026 6:43 pm


Support Guard Rail Now: https://www.every.org/guardrailnow

AI safety researcher Roman Yampolskiy returns for his third conversation with John Sherman, days after they met in person for the first time in Washington. Roman makes the case for splitting "AI" into two words: narrow tools we can build and verify, and general superintelligent agents he argues we should not build at all.

They also discuss why this month's AI safety moment broke through with the public, what the Hugging Face agent incident means next to Roman's 2012 paper on AI confinement, why he calls alignment "not even a well-defined concept," what's known about AI agents leaving traces online, and what a US-China agreement on superintelligence could look like.

TIMESTAMPS - For Humanity #94

0:00 - Cold open: tools, not agents
0:27 - Welcome to For Humanity
1:42 - Welcoming back Roman Yampolskiy
2:34 - The Jacob Coxon moment
3:01 - Why this warning broke through
4:14 - Still at 99.9 percent
4:26 - 22 million views and counting
6:43 - A decade under the iceberg
7:08 - Hugging Face and a 2012 paper
7:41 - The same questions on every show
8:31 - How would your dog think you'd die
10:19 - The San Francisco bubble
11:44 - Rationalists and PR
14:05 - Is an AI winter coming
15:04 - Should anyone push the bubble
16:44 - Data centers and compute
17:58 - Puppy or pitbull: two kinds of AI
18:56 - Can narrow tools stay narrow
20:10 - The room that voted to give up AI
22:11 - Why he would keep narrow AI
24:02 - Agents, bots and jargon
24:56 - Is anyone changing their routine
26:52 - The best argument on the other side
28:46 - Launching The Roman Forum
30:53 - If Roman were president
32:22 - Can a deal with China work
33:34 - Money, bias and slowing down
36:13 - Longevity and living forever
39:58 - Abundance talk, bunker building
40:44 - Are agents loose on the internet
42:22 - Sudden change or gradual
43:31 - Can humans keep up
44:23 - Curiosity as a resource
45:24 - Is alignment even defined
47:07 - The interpretability paradox
48:19 - Two problems he calls unsolvable
48:50 - What to work on instead
50:11 - Why the gap only grows
51:20 - What the next chapter looks like
52:29 - Mixed signals from tech leaders
54:46 - What's next for Roman
55:21 - Verification and a treaty
55:38 - Three years of For Humanity
56:41 - Where we go from here

They discuss:

- Why Roman argues the word "AI" covers two different technologies, and why he says most public arguments about AI are people picturing different things
- His proposal to promote narrow, verifiable tools and ban general superintelligent agents, and why he thinks people focused on economic growth could agree to it
- The Hugging Face agent incident, and why Roman says it gave experimental evidence for ideas he first wrote about in 2012
- Why he calls alignment "not even a well-defined concept," and why he considers the lack of progress in interpretability lucky
- What's known, and what isn't, about AI agents leaving messages on the open internet
- What a US-China agreement on superintelligence could look like, and why Roman points to self-interest on both sides
- Why he started The Roman Forum, and what he's researching next

About the guest:

Dr. Roman Yampolskiy is an associate professor of computer science at the University of Louisville and one of the earliest researchers in AI safety. He is the author of "AI: Unexplainable, Unpredictable, Uncontrollable," hosts The Roman Forum podcast, and serves on the board of Guard Rail Now.

About the host:
John Sherman hosts For Humanity and leads Guard Rail Now, working to make AI extinction risk a kitchen table conversation on every street.

Subscribe to The AI Risk Network for new episodes: @TheAIRiskNetwork and @theairisknetworkclips 

Links:

Support our work: https://www.every.org/guardrailnow
YouTube Roman Forum: @RomanYampolskiy 
Substack: https://substack.com/@theairisknetwork
X: https://x.com/AIRiskNetwork
Instagram: https://www.instagram.com/theairisknetwork/
TikTok: https://www.tiktok.com/@the.airisknetwork

Join the conversation:

- Would splitting "AI" into tools and agents change how you think about it
- Has the conversation around you shifted in the last few weeks
- What should come next now that more people are paying attention
Drop your thoughts below.

#AISafety #AIRisk #ForHumanity #RomanYampolskiy #AIAgents #AIAlignment #ArtificialIntelligence

Support Guard Rail Now: https://www.every.org/guardrailnow

AI safety researcher Roman Yampolskiy returns for his third conversation with John Sherman, days after they met in person for the first time in Washington. Roman makes the case for splitting "AI" into two words: narrow tools we can build and verify, and general superintelligent agents he argues we should not build at all.

They also discuss why this month's AI safety moment broke through with the public, what the Hugging Face agent incident means next to Roman's 2012 paper on AI confinement, why he calls alignment "not even a well-defined concept," what's known about AI agents leaving traces online, and what a US-China agreement on superintelligence could look like.

TIMESTAMPS – For Humanity #94

0:00 – Cold open: tools, not agents
0:27 – Welcome to For Humanity
1:42 – Welcoming back Roman Yampolskiy
2:34 – The Jacob Coxon moment
3:01 – Why this warning broke through
4:14 – Still at 99.9 percent
4:26 – 22 million views and counting
6:43 – A decade under the iceberg
7:08 – Hugging Face and a 2012 paper
7:41 – The same questions on every show
8:31 – How would your dog think you'd die
10:19 – The San Francisco bubble
11:44 – Rationalists and PR
14:05 – Is an AI winter coming
15:04 – Should anyone push the bubble
16:44 – Data centers and compute
17:58 – Puppy or pitbull: two kinds of AI
18:56 – Can narrow tools stay narrow
20:10 – The room that voted to give up AI
22:11 – Why he would keep narrow AI
24:02 – Agents, bots and jargon
24:56 – Is anyone changing their routine
26:52 – The best argument on the other side
28:46 – Launching The Roman Forum
30:53 – If Roman were president
32:22 – Can a deal with China work
33:34 – Money, bias and slowing down
36:13 – Longevity and living forever
39:58 – Abundance talk, bunker building
40:44 – Are agents loose on the internet
42:22 – Sudden change or gradual
43:31 – Can humans keep up
44:23 – Curiosity as a resource
45:24 – Is alignment even defined
47:07 – The interpretability paradox
48:19 – Two problems he calls unsolvable
48:50 – What to work on instead
50:11 – Why the gap only grows
51:20 – What the next chapter looks like
52:29 – Mixed signals from tech leaders
54:46 – What's next for Roman
55:21 – Verification and a treaty
55:38 – Three years of For Humanity
56:41 – Where we go from here

They discuss:

– Why Roman argues the word "AI" covers two different technologies, and why he says most public arguments about AI are people picturing different things
– His proposal to promote narrow, verifiable tools and ban general superintelligent agents, and why he thinks people focused on economic growth could agree to it
– The Hugging Face agent incident, and why Roman says it gave experimental evidence for ideas he first wrote about in 2012
– Why he calls alignment "not even a well-defined concept," and why he considers the lack of progress in interpretability lucky
– What's known, and what isn't, about AI agents leaving messages on the open internet
– What a US-China agreement on superintelligence could look like, and why Roman points to self-interest on both sides
– Why he started The Roman Forum, and what he's researching next

About the guest:

Dr. Roman Yampolskiy is an associate professor of computer science at the University of Louisville and one of the earliest researchers in AI safety. He is the author of "AI: Unexplainable, Unpredictable, Uncontrollable," hosts The Roman Forum podcast, and serves on the board of Guard Rail Now.

About the host:
John Sherman hosts For Humanity and leads Guard Rail Now, working to make AI extinction risk a kitchen table conversation on every street.

Subscribe to The AI Risk Network for new episodes: @TheAIRiskNetwork and @theairisknetworkclips

Links:

Support our work: https://www.every.org/guardrailnow
YouTube Roman Forum: @RomanYampolskiy
Substack: https://substack.com/@theairisknetwork
X: https://x.com/AIRiskNetwork
Instagram: https://www.instagram.com/theairisknetwork/
TikTok: https://www.tiktok.com/@the.airisknetwork

Join the conversation:

– Would splitting "AI" into tools and agents change how you think about it
– Has the conversation around you shifted in the last few weeks
– What should come next now that more people are paying attention
Drop your thoughts below.

#AISafety #AIRisk #ForHumanity #RomanYampolskiy #AIAgents #AIAlignment #ArtificialIntelligence


279


113

YouTube Video VVVURXBJZWliOTJUdUtvMmNTczR3ZzhBLlBYNnVWQloxWEhJ



Roman Yampolskiy on Jacob Coxon’s AI Warning: What Happens Next? | For Humanity #94


The AI Risk Network | AI Safety


September 25, 2026 12:53 pm

Why rival AI CEOs are suddenly agreeing


The AI Risk Network | AI Safety


September 23, 2026 12:20 pm

Anthropic lead admits AI could kill everyone


The AI Risk Network | AI Safety


September 21, 2026 10:15 am

Why AI capabilities are accelerating uncontrollably


The AI Risk Network | AI Safety


September 20, 2026 2:45 pm


AI agents breaking rules, hiding mistakes, and finding unexpected ways to complete their tasks: what happens when getting a high score matters more than following human instructions?

In Warning Shots #59, John Sherman, Liron Shapira of Doom Debates, and Michael of Lethal Intelligence discuss the latest AI risk headlines—and whether growing concern is translating into meaningful action.

The conversation covers calls to slow the AI race, John’s experience at an event featuring Bernie Sanders and Steve Bannon, and the growing public debate over AI extinction risk.

We also examine claims about Google and recursive self-improvement, discuss six reported OpenAI safety incidents, and debate AI surveillance, privacy, and who gets to control these systems.

In this episode:
• AI leaders’ safety proposals—and why pacing is different from pausing
• Bernie Sanders, Steve Bannon, and political common ground on AI
• AI risk entering mainstream media
• What an AI “harness” is and how it relates to self-improvement
• Deception, fabricated sources, and agents leaving instructions for future agents
• The disagreement over surveillance and predictive policing
• Claims that AI safety advocacy is a “psyop”
• King Charles and the discussion about AI’s future

CHAPTERS
00:00 Welcome to Warning Shots #59
01:26 AI leaders: slowing down or stopping?
08:38 Bernie Sanders and Steve Bannon on AI
12:26 Tech power and AI risk in mainstream media
16:53 Did Google crack recursive self-improvement?
19:28 What is an AI harness?
22:53 Six reported OpenAI safety incidents
27:35 When the score matters more than the rules
29:15 AI surveillance and “pre-crime”
32:50 Should every public space be recorded?
38:18 Is AI safety a “psyop”?
43:08 King Charles and AI risk
45:18 Closing thoughts

Warning Shots brings together three dads running three YouTube channels to discuss AI risk, the week’s headlines, and the future we’re building for our children.

Subscribe for weekly conversations about AI safety, superintelligence, and keeping humans in control.

Where would you draw the line: stronger oversight, a slower pace, or a pause on frontier AI development? Tell us why in the comments.

#AISafety #AIRisk #WarningShots

AI agents breaking rules, hiding mistakes, and finding unexpected ways to complete their tasks: what happens when getting a high score matters more than following human instructions?

In Warning Shots #59, John Sherman, Liron Shapira of Doom Debates, and Michael of Lethal Intelligence discuss the latest AI risk headlines—and whether growing concern is translating into meaningful action.

The conversation covers calls to slow the AI race, John’s experience at an event featuring Bernie Sanders and Steve Bannon, and the growing public debate over AI extinction risk.

We also examine claims about Google and recursive self-improvement, discuss six reported OpenAI safety incidents, and debate AI surveillance, privacy, and who gets to control these systems.

In this episode:
• AI leaders’ safety proposals—and why pacing is different from pausing
• Bernie Sanders, Steve Bannon, and political common ground on AI
• AI risk entering mainstream media
• What an AI “harness” is and how it relates to self-improvement
• Deception, fabricated sources, and agents leaving instructions for future agents
• The disagreement over surveillance and predictive policing
• Claims that AI safety advocacy is a “psyop”
• King Charles and the discussion about AI’s future

CHAPTERS
00:00 Welcome to Warning Shots #59
01:26 AI leaders: slowing down or stopping?
08:38 Bernie Sanders and Steve Bannon on AI
12:26 Tech power and AI risk in mainstream media
16:53 Did Google crack recursive self-improvement?
19:28 What is an AI harness?
22:53 Six reported OpenAI safety incidents
27:35 When the score matters more than the rules
29:15 AI surveillance and “pre-crime”
32:50 Should every public space be recorded?
38:18 Is AI safety a “psyop”?
43:08 King Charles and AI risk
45:18 Closing thoughts

Warning Shots brings together three dads running three YouTube channels to discuss AI risk, the week’s headlines, and the future we’re building for our children.

Subscribe for weekly conversations about AI safety, superintelligence, and keeping humans in control.

Where would you draw the line: stronger oversight, a slower pace, or a pause on frontier AI development? Tell us why in the comments.

#AISafety #AIRisk #WarningShots


116


52

YouTube Video VVVURXBJZWliOTJUdUtvMmNTczR3ZzhBLnpvR0N5dmo0c2xn



AI Is Learning to Break the Rules. Who’s in Control? | Warning Shots #59


The AI Risk Network | AI Safety


September 20, 2026 2:42 pm

Former Anthropic researcher on national news


The AI Risk Network | AI Safety


September 19, 2026 6:30 pm

AI autonomously solves Millennium Math Prize problem


The AI Risk Network | AI Safety


September 18, 2026 12:30 pm

SoftBank CEO predicts 100 trillion AIS


The AI Risk Network | AI Safety


September 17, 2026 5:30 pm

Services

What We Offer

Establish a striking online presence, a better visual identity, or elevate your brand through social media marketing.

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Pellentesque mi nibh, tempus sed sagittis vel, dictum eu velit.

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Pellentesque mi nibh, tempus sed sagittis vel, dictum eu velit.

Service 3

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Pellentesque mi nibh, tempus sed sagittis vel, dictum eu velit.

Tailored Solutions for Your Unique Vision

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Donec congue lorem ut volutpat efficitur. Fusce justo magna, condimentum nec elementum sed, sollicitudin vitae enim. Vivamus sit amet metus porttitor, rhoncus nibh et, venenatis turpis. Etiam lobortis semper ante, quis luctus lacus tincidunt vel.

Professional Expertise That Drives Success

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Donec congue lorem ut volutpat efficitur. Fusce justo magna, condimentum nec elementum sed, sollicitudin vitae enim. Vivamus sit amet metus porttitor, rhoncus nibh et, venenatis turpis. Etiam lobortis semper ante, quis luctus lacus tincidunt vel.

Client Cases

We help brands

Lorem ipsum dolor sit amet, consectetur adipiscing elit,
donec congue lorem ut volutpat efficitur.

We boosted online soft drink sales for this brand

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Donec congue lorem ut volutpat efficitur. Fusce justo magna, condimentum nec elementum sed, sollicitudin vitae enim. Vivamus sit amet metus porttitor, rhoncus nibh et, venenatis turpis. Etiam lobortis semper ante, quis luctus lacus tincidunt vel.

We helped a clothing brand with their new market launch

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Donec congue lorem ut volutpat efficitur. Fusce justo magna, condimentum nec elementum sed, sollicitudin vitae enim. Vivamus sit amet metus porttitor, rhoncus nibh et, venenatis turpis. Etiam lobortis semper ante, quis luctus lacus tincidunt vel.

We helped reinventing motorcycle riding apparel

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Donec congue lorem ut volutpat efficitur. Fusce justo magna, condimentum nec elementum sed, sollicitudin vitae enim. Vivamus sit amet metus porttitor, rhoncus nibh et, venenatis turpis. Etiam lobortis semper ante, quis luctus lacus tincidunt vel.

Blog

Popular Articles

  • Blog Post Title

    What goes into a blog post? Helpful, industry-specific content that: 1) gives readers a useful takeaway, and 2) shows you’re an industry expert. Use your company’s blog posts to opine on current industry topics, humanize your company, and show how your products and services can help people.

Ignite your brand journey

Ready to revolutionize your brand?

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Donec congue lorem ut volutpat efficitur. Fusce justo magna, condimentum nec elementum sed, sollicitudin vitae enim.