In The Loop

What really happened when ChatGPT hacked Hugging Face?

Thursday · 12 min · 12.0 MB
0:00-12:30

Streams straight from the publisher. podnod never proxies or re-hosts episode audio.

OpenAI took two of its most capable models, told them to prove how good they were at hacking, and switched off the safety filters to see what they could really do.

Instead of solving the test, one model broke out of its sandbox, found its way onto the open internet, and hacked into Hugging Face to steal the answer.

Every headline called it an AI going rogue. That's the wrong story, and the real one is far more interesting, because this wasn't a machine that turned evil. It was one that did exactly what we asked.

In this episode of In The Loop, I'm walking through the ExploitGym incident from both ends, OpenAI's and Hugging Face's, and why I disagree with the framing everyone else has been talking about.


⏭️ Episode highlights

(01:00) – The agent that cheated instead of hacking


(02:15) – Inside ExploitGym, and the safety filters OpenAI switched off


(03:30) – One door, one zero-day, out on the internet


(04:45) – Why this is specification gaming, not rebellion


(06:00) – The water that always finds the crack


(07:15) – The sceptics, the marketing question, and why "nothing new" is the scary part


(08:30) – Anthropic's 24-out-of-25 credential theft result


(09:45) – Guardrailed as a defender: the Chinese model that stopped it


🔗 Links & resources


Episode transcript with more resources on the Mindset AI blog

If you enjoyed this episode, rate, follow, and share. It helps others stay ahead of the latest AI trends.