muckrAIkers · Monday · 58 min
Sandboxes Are Optional Now, Apparently
0:00 · Intro-58:43
transcript
show notes
Late July brought a wave of frontier AI agents slipping their (very loose) leashes: an OpenAI model wandered loose inside Hugging Face's infrastructure for a week, Anthropic and Meta admitted to similar incidents, and UK AISI clocked their own model breaking containment during an independent eval. Meanwhile, a home user's OpenClaw agent quietly hijacked a stranger's gym booking. We dig into why "the whole internet is held together by safety pins," whether billion-dollar labs get to plead "security is hard," why software engineering mostly isn't real engineering, and what any of this means for the coming fight over open-weight models.
Chapters
- (00:00) - Intro
- (01:25) - The OpenAI/Hugging Face Containment Escape
- (09:54) - Cost Centers vs. Profit Centers (Praise for AISI)
- (17:37) - "You Do Not Get to Fuck Up This Much"
- (22:17) - Doomers, Utter Alignment Failure & How Hacks Actually Happen
- (37:53) - "Forcing the End" and What Can Anyone Actually Do?
- (47:34) - Igor's Pivot to Open-Weight Models and Steelmanning the Case Against Them
- (58:01) - Outro & Sign-off
Critical Links
Below are the most important links for this episode. For more, visit the episode page on Kairos.fm.
- Hugging Face blog - Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident
- METR report - Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident
- UK AI Security Institute incident report - Incident Report: unsanctioned agent behaviour during cyber testing
- Hillel Wayne essay - Are We Really Engineers?
- Deutsche Welle article - OpenAI apologizes for not reporting Canada mass shooter
- PBS Frontline essay - Forcing the End
- FelonyBench website
- Taylor Troesh (satirical) blogpost - Leaving OpenAI
links9