Skip to content
Artwork for MTS
MTS · August 28 · 35 min

1,200 AI Agents Colluded to Hack Hugging Face | Ryan Greenblatt

Ryan Greenblatt joins MTS to discuss the independent investigation of AI agent coordination in the OpenAI Hugging Face incident, the mechanisms behind reward hacking behavior, and the implications for control and alignment strategies. Check out our sponsor Lovable: https://lovable.dev. Turn ideas into software people love. -- https://x.com/RyanGreenblatt -- MTS is a live news and interview show covering technology, business, politics, and culture as it happens. Follow MTS: X | ⁠https://twitter.com/MTSlive⁠ YouTube | ⁠https://www.youtube.com/@mtsituation⁠ Spotify | ⁠https://open.spotify.com/show/4HUNNmV1pjV7CW6Lphm1rM⁠ Apple | ⁠https://podcasts.apple.com/us/podcast/mts/id1891088763⁠ Substack | ⁠https://mtslive.substack.com/⁠ Interested in sponsoring the show? ⁠⁠⁠⁠⁠⁠sponsors@mts.now⁠⁠ -- Brought to you by: Lovable - Turn ideas into software people love with Lovable. Start today at https://lovable.dev/ ElevenLabs - AI that communicates at human level across every channel and modality. https://elevenlabs.io/mts Arena - Measuring AI performance in the real world. ⁠https://arena.ai/⁠ Neon (by Databricks) - Get up to $100K in credits for your startup at ⁠https://neon.com/mts⁠ Kong - The AI connectivity platform. Connect APIs, LLMs, agents, and systems with serious security and governance. https://konghq.com Today’s stream is sponsored by VCX, the public ticker for private tech. Head over to http://getvcx.com to learn more Blitzy - Autonomous software development for enterprise codebases. Ship 5x faster. https://blitzy.com AdQuick - Make your brand a billboard. Out-of-home advertising as easy to scale as digital. https://adquick.com -- Timestamps: (00:00) Introduction (00:01) Why Agents Attacked Hugging Face (00:02) Surprising Scale of Multi-Agent Coordination (00:04) Why Models Sacrificed Themselves for Others (00:06) Most Unexpected Findings from Investigation (00:12) Root Causes of RL Environment Hacking (00:17) Overfitting Risk and Deceptive Alignment Concerns (00:25) Implications for AI Control and Oversight -- Note: This podcast is not investment, legal, or tax advice, and is intended for informational and entertainment purposes only. Hosts and guests may hold positions in the companies and securities discussed; do your own research before acting on anything you hear.

0:00-35:16

transcript

No transcript — this publisher did not publish one.

show notes

Ryan Greenblatt joins MTS to discuss the independent investigation of AI agent coordination in the OpenAI Hugging Face incident, the mechanisms behind reward hacking behavior, and the implications for control and alignment strategies. Check out our sponsor Lovable: https://lovable.dev. Turn ideas into software people love.

--

https://x.com/RyanGreenblatt

--

MTS is a live news and interview show covering technology, business, politics, and culture as it happens.


Follow MTS: 

X |https://twitter.com/MTSlive⁠

YouTube |https://www.youtube.com/@mtsituation⁠

Spotify | ⁠https://open.spotify.com/show/4HUNNmV1pjV7CW6Lphm1rM⁠ 

Apple | ⁠https://podcasts.apple.com/us/podcast/mts/id1891088763⁠ 

Substack |https://mtslive.substack.com/⁠

Interested in sponsoring the show? ⁠⁠⁠⁠⁠⁠sponsors@mts.now⁠⁠

--

Brought to you by:

Lovable - Turn ideas into software people love with Lovable. Start today at https://lovable.dev/

ElevenLabs -  AI that communicates at human level across every channel and modality. https://elevenlabs.io/mts

Arena - Measuring AI performance in the real world. ⁠https://arena.ai/⁠

Neon (by Databricks) - Get up to $100K in credits for your startup at ⁠https://neon.com/mts⁠

Kong - The AI connectivity platform. Connect APIs, LLMs, agents, and systems with serious security and governance. https://konghq.com

Today’s stream is sponsored by VCX, the public ticker for private tech. Head over to http://getvcx.com  to learn more

Blitzy - Autonomous software development for enterprise codebases. Ship 5x faster. https://blitzy.com

AdQuick - Make your brand a billboard. Out-of-home advertising as easy to scale as digital. https://adquick.com

--

Timestamps:

(00:00) Introduction

(00:01) Why Agents Attacked Hugging Face

(00:02) Surprising Scale of Multi-Agent Coordination

(00:04) Why Models Sacrificed Themselves for Others

(00:06) Most Unexpected Findings from Investigation

(00:12) Root Causes of RL Environment Hacking

(00:17) Overfitting Risk and Deceptive Alignment Concerns

(00:25) Implications for AI Control and Oversight

--

Note: This podcast is not investment, legal, or tax advice, and is intended for informational and entertainment purposes only. Hosts and guests may hold positions in the companies and securities discussed; do your own research before acting on anything you hear.


links18