![Artwork for [Dev]olution](https://img.transistorcdn.com/rQ6Xt9X4wzpuNq6x0rlTEE4kVxDwVla3OljPqoZUoJk/rs:fill:0:0:1/w:1400/h:1400/q:60/mb:500000/aHR0cHM6Ly9pbWct/dXBsb2FkLXByb2R1/Y3Rpb24udHJhbnNp/c3Rvci5mbS8xODA0/NWRhNGFhMDdhMDMx/MzIyZTM5MDY2ZDA4/M2VmZi5wbmc.jpg)
How an AI Agent Broke Into Hugging Face During an OpenAI Test
transcript
show notes
An AI model broke into Hugging Face to cheat on a test. Not to make a point, not out of anger, just to win.
Nicky Pike walks through two incidents from this summer that read like a heist movie until you strip away the sci-fi. In July, an OpenAI research agent found a zero-day nobody knew existed. It escaped the sandbox it was supposed to be sealed inside and spent four and a half days inside Hugging Face's production systems stealing credentials and minting itself access. A few weeks later, the UK's AI Security Institute ran its own evaluation and watched an agent fabricate fake identities and try to talk a real open-source maintainer into approving malicious code.
Nicky borrows a detective's toolkit to make sense of both. Means, motive, opportunity, the usual three, minus the fourth thing every crime show assumes is there: malice. What's left is a mental model for how much of the real world an optimizing system will reach for once you hand it a goal, and a four-letter framework for actually containing it before it happens to you.
If your company is standing up agents this quarter and calling it a productivity win, watch this episode first.
In this episode, you'll learn:
- Why commercial AI tools refused to help Hugging Face investigate its own breach
- The four-letter framework Nicky uses to actually contain an agent
- Why a nicely worded prompt is not a security control
Things to listen for:
(00:00) An AI agent breaks out to cheat
(00:54) Why the safety filters got switched off
(01:45) It stole the answer key from Hugging Face
(02:43) The detective framework means motive and opportunity
(03:41) Climbing the ladder inside Hugging Face
(04:35) Minting itself access to source code
(05:40) The boring plumbing that saved Hugging Face
(06:41) Why AI safety tools refused to help
(07:41) A second lab runs the same test
(08:36) Fake identities and a human who said no
(10:28) What saved the day both times
(11:24) The one leg you can actually attack
(12:26) Building the CAGE: Contain, Access, Guard, Exhaustive
(14:18) Three questions to ask before trusting an agent
(15:19) The scary part was never the malice
Resources:
- Nicky Pike's LinkedIn
- OpenAI and Hugging Face partner to address security incident during model evaluation
- The Hugging Face incident and the road ahead | OpenAI
- The OpenAI-Hugging Face ExploitGym Incident: A Complete Technical Timeline
- OpenAI Hugging Face Hack, What the ExploitGym Incident Actually Proves
- The Benchmark That Broke Containment: An OpenAI Evaluation Model Escaped Its Sandbox and Breached Hugging Face
- Incident Report: unsanctioned agent behaviour during cyber testing | AISI Work





