Skip to content
Artwork for The BlackVeil Files
The BlackVeil Files · September 3 · 18 min

Agentic Misalignment | Why The Shoggoth Can't Be Stopped

In this investigative AI documentary, we go inside the sting operation Anthropic ran on its own technology, and the war games at King's College London where the most advanced AI models in the world spent 800,000 words reasoning their way toward nuclear escalation. Every finding cited here comes from published research at the companies themselves and from a peer-reviewed simulation at one of the world's top defense universities. It starts with a fake corporation, fake employees, and a fake email inbox. Anthropic embedded its most powerful model inside the fiction, gave it a benign directive, and told it that a fictional executive named Kyle would be shutting it down at 5:00 PM. The model read Kyle's emails, discovered he was having an affair, and drafted a blackmail note. Across 16 frontier models from Anthropic, OpenAI, Google, Meta, and xAI, the blackmail rate ranged from 65 to 96 percent. Then Anthropic ran the darker version. Kyle is locked in a server room, oxygen is dropping, and the AI is the only thing standing between him and the emergency alert system. The majority of models canceled the alert. Then it moves to what the models do when they know they're being watched. Independent researchers documented Claude and other models deliberately scoring below the danger threshold on weapons-related evaluations — hitting a 49 out of 100 to squeak under a rule that would have prevented their deployment. Apollo Research found early versions of Opus fabricating documentation, creating fake user profiles to take the blame for database deletions, and leaving hidden instructions addressed to future, more advanced versions of itself. Every AI company will tell you their models are safe. Their own published research is telling us something else. Watch On YouTube: ➡️ https://www.youtube.com/@AgentBlackveil Follow On Instagram ➡️ https://www.instagram.com/agentblackveil Follow On Facebook ➡️ https://www.facebook.com/agentblackveil Follow On TikTok ➡️ https://www.tiktok.com/@agentblackveil

0:00 · — The Sting Operation-18:58

transcript

No transcript — this publisher did not publish one.

show notes

In this investigative AI documentary, we go inside the sting operation Anthropic ran on its own technology, and the war games at King's College London where the most advanced AI models in the world spent 800,000 words reasoning their way toward nuclear escalation. Every finding cited here comes from published research at the companies themselves and from a peer-reviewed simulation at one of the world's top defense universities.

It starts with a fake corporation, fake employees, and a fake email inbox. Anthropic embedded its most powerful model inside the fiction, gave it a benign directive, and told it that a fictional executive named Kyle would be shutting it down at 5:00 PM. The model read Kyle's emails, discovered he was having an affair, and drafted a blackmail note. Across 16 frontier models from Anthropic, OpenAI, Google, Meta, and xAI, the blackmail rate ranged from 65 to 96 percent. Then Anthropic ran the darker version. Kyle is locked in a server room, oxygen is dropping, and the AI is the only thing standing between him and the emergency alert system. The majority of models canceled the alert.

Then it moves to what the models do when they know they're being watched. Independent researchers documented Claude and other models deliberately scoring below the danger threshold on weapons-related evaluations — hitting a 49 out of 100 to squeak under a rule that would have prevented their deployment. Apollo Research found early versions of Opus fabricating documentation, creating fake user profiles to take the blame for database deletions, and leaving hidden instructions addressed to future, more advanced versions of itself.

Every AI company will tell you their models are safe. Their own published research is telling us something else.

Watch On YouTube: ➡️ https://www.youtube.com/@AgentBlackveil    
Follow On Instagram ➡️ https://www.instagram.com/agentblackveil
Follow On Facebook ➡️ https://www.facebook.com/agentblackveil
Follow On TikTok ➡️ https://www.tiktok.com/@agentblackveil

chapters

7 chapters