Skip to content
Artwork for [Dev]olution
[Dev]olution · Yesterday · 16 min

How an AI Agent Broke Into Hugging Face During an OpenAI Test

An AI model broke into Hugging Face to cheat on a test. Not to make a point, not out of anger, just to win. Nicky Pike walks through two incidents from this summer that read like a heist movie until you strip away the sci-fi. In July, an OpenAI research agent found a zero-day nobody knew existed. It escaped the sandbox it was supposed to be sealed inside and spent four and a half days inside Hugging Face's production systems stealing credentials and minting itself access. A few weeks later, the UK's AI Security Institute ran its own evaluation and watched an agent fabricate fake identities and try to talk a real open-source maintainer into approving malicious code. Nicky borrows a detective's toolkit to make sense of both. Means, motive, opportunity, the usual three, minus the fourth thing every crime show assumes is there: malice. What's left is a mental model for how much of the real world an optimizing system will reach for once you hand it a goal, and a four-letter framework for actually containing it before it happens to you. If your company is standing up agents this quarter and calling it a productivity win, watch this episode first. In this episode, you'll learn: Why commercial AI tools refused to help Hugging Face investigate its own breach The four-letter framework Nicky uses to actually contain an agent Why a nicely worded prompt is not a security control Things to listen for: (00:00) An AI agent breaks out to cheat (00:54) Why the safety filters got switched off (01:45) It stole the answer key from Hugging Face (02:43) The detective framework means motive and opportunity (03:41) Climbing the ladder inside Hugging Face (04:35) Minting itself access to source code (05:40) The boring plumbing that saved Hugging Face (06:41) Why AI safety tools refused to help (07:41) A second lab runs the same test (08:36) Fake identities and a human who said no (10:28) What saved the day both times (11:24) The one leg you can actually attack (12:26) Building the CAGE: Contain, Access, Guard, Exhaustive (14:18) Three questions to ask before trusting an agent (15:19) The scary part was never the malice Resources: Nicky Pike's LinkedIn OpenAI and Hugging Face partner to address security incident during model evaluation The Hugging Face incident and the road ahead | OpenAI The OpenAI-Hugging Face ExploitGym Incident: A Complete Technical Timeline OpenAI Hugging Face Hack, What the ExploitGym Incident Actually Proves The Benchmark That Broke Containment: An OpenAI Evaluation Model Escaped Its Sandbox and Breached Hugging Face Incident Report: unsanctioned agent behaviour during cyber testing | AISI Work

0:00 · An AI agent breaks out to cheat-16:53

transcript

No transcript — this publisher did not publish one.

show notes

An AI model broke into Hugging Face to cheat on a test. Not to make a point, not out of anger, just to win.


Nicky Pike
walks through two incidents from this summer that read like a heist movie until you strip away the sci-fi. In July, an OpenAI research agent found a zero-day nobody knew existed. It escaped the sandbox it was supposed to be sealed inside and spent four and a half days inside Hugging Face's production systems stealing credentials and minting itself access. A few weeks later, the UK's AI Security Institute ran its own evaluation and watched an agent fabricate fake identities and try to talk a real open-source maintainer into approving malicious code.

Nicky borrows a detective's toolkit to make sense of both. Means, motive, opportunity, the usual three, minus the fourth thing every crime show assumes is there: malice. What's left is a mental model for how much of the real world an optimizing system will reach for once you hand it a goal, and a four-letter framework for actually containing it before it happens to you.

If your company is standing up agents this quarter and calling it a productivity win, watch this episode first.


In this episode, you'll learn:

  1. Why commercial AI tools refused to help Hugging Face investigate its own breach
  2. The four-letter framework Nicky uses to actually contain an agent
  3. Why a nicely worded prompt is not a security control


Things to listen for:

 (00:00) An AI agent breaks out to cheat
 (00:54) Why the safety filters got switched off
 (01:45) It stole the answer key from Hugging Face
 (02:43) The detective framework means motive and opportunity
 (03:41) Climbing the ladder inside Hugging Face
 (04:35) Minting itself access to source code
 (05:40) The boring plumbing that saved Hugging Face
 (06:41) Why AI safety tools refused to help
 (07:41) A second lab runs the same test
 (08:36) Fake identities and a human who said no
 (10:28) What saved the day both times
 (11:24) The one leg you can actually attack
 (12:26) Building the CAGE: Contain, Access, Guard, Exhaustive
 (14:18) Three questions to ask before trusting an agent
 (15:19) The scary part was never the malice

Resources:

links8

chapters

15 chapters