Skip to content
Artwork for muckrAIkers
muckrAIkers · Monday · 58 min

Sandboxes Are Optional Now, Apparently

Late July brought a wave of frontier AI agents slipping their (very loose) leashes: an OpenAI model wandered loose inside Hugging Face's infrastructure for a week, Anthropic and Meta admitted to similar incidents, and UK AISI clocked their own model breaking containment during an independent eval. Meanwhile, a home user's OpenClaw agent quietly hijacked a stranger's gym booking. We dig into why "the whole internet is held together by safety pins," whether billion-dollar labs get to plead "security is hard," why software engineering mostly isn't real engineering, and what any of this means for the coming fight over open-weight models. Chapters (00:00) - Intro (01:25) - The OpenAI/Hugging Face Containment Escape (09:54) - Cost Centers vs. Profit Centers (Praise for AISI) (17:37) - "You Do Not Get to Fuck Up This Much" (22:17) - Doomers, Utter Alignment Failure & How Hacks Actually Happen (37:53) - "Forcing the End" and What Can Anyone Actually Do? (47:34) - Igor's Pivot to Open-Weight Models and Steelmanning the Case Against Them (58:01) - Outro & Sign-off Critical Links Below are the most important links for this episode. For more, visit the episode page on Kairos.fm. Hugging Face blog - Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident METR report - Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident UK AI Security Institute incident report - Incident Report: unsanctioned agent behaviour during cyber testing Hillel Wayne essay - Are We Really Engineers? Deutsche Welle article - OpenAI apologizes for not reporting Canada mass shooter PBS Frontline essay - Forcing the End FelonyBench website Taylor Troesh (satirical) blogpost - Leaving OpenAI

0:00 · Intro-58:43

transcript

No transcript — this publisher did not publish one.

show notes

Late July brought a wave of frontier AI agents slipping their (very loose) leashes: an OpenAI model wandered loose inside Hugging Face's infrastructure for a week, Anthropic and Meta admitted to similar incidents, and UK AISI clocked their own model breaking containment during an independent eval. Meanwhile, a home user's OpenClaw agent quietly hijacked a stranger's gym booking. We dig into why "the whole internet is held together by safety pins," whether billion-dollar labs get to plead "security is hard," why software engineering mostly isn't real engineering, and what any of this means for the coming fight over open-weight models.

Chapters

  • (00:00) - Intro
  • (01:25) - The OpenAI/Hugging Face Containment Escape
  • (09:54) - Cost Centers vs. Profit Centers (Praise for AISI)
  • (17:37) - "You Do Not Get to Fuck Up This Much"
  • (22:17) - Doomers, Utter Alignment Failure & How Hacks Actually Happen
  • (37:53) - "Forcing the End" and What Can Anyone Actually Do?
  • (47:34) - Igor's Pivot to Open-Weight Models and Steelmanning the Case Against Them
  • (58:01) - Outro & Sign-off

Critical Links
Below are the most important links for this episode. For more, visit the episode page on Kairos.fm.
  • Hugging Face blog - Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident
  • METR report - Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident
  • UK AI Security Institute incident report - Incident Report: unsanctioned agent behaviour during cyber testing
  • Hillel Wayne essay - Are We Really Engineers?
  • Deutsche Welle article - OpenAI apologizes for not reporting Canada mass shooter
  • PBS Frontline essay - Forcing the End
  • FelonyBench website
  • Taylor Troesh (satirical) blogpost - Leaving OpenAI
links9

chapters

8 chapters