“Anthropic Has Some Alignment Problems” by Zvi
transcript
show notes
Oh, good. They noticed.
Anthropic, too, is planning to bring METR inside for an independent review of their own incidents, where three times a Claude model started hacking outside things during an eval, and where Mythos 5 did various ‘unauthorized actions,’ by which we mean tried to hack various real-world things, during a UK AISI cybersecurity eval.
Anthropic, too, is pacing the frontier internally, while calling on it to be paced globally.
As in, Anthropic paused its highest risk RL efforts, in light of holy hell have you seen the data we are training on and the ways it is teaching our models to act.
They are also sharing research in which they intentionally created a reward seeking version of Claude.
Scheduling note: Fable 5.1 has been released. I will aim to cover that starting Friday. OpenAI is also planning to release Astra soon, which I would cover after Fable.
Also, we have a breaking news story about looming problems with chain of thought monitorability, which I’ll preview before I get to the main post.
Table of Contents
- This Just In.
- Anthropic Parallel Pauses.
- Pause The Data Brokers. [...]
---
Outline:
(01:16) This Just In
(02:43) Anthropic Parallel Pauses
(08:22) Pause The Data Brokers
(09:54) Pacing the Frontier
(11:54) Misalignment Assessment
(13:39) Defects In Training Environments Disproportionately Cause Cheating
(14:59) Creating Reward Hacker Opus
(19:33) Undo It
(21:00) Mistakes Were Made
(23:33) Internal Security Posture
(25:26) One Does Not Simply Fix The RL Environments
---
First published:
September 2nd, 2026
Source:
https://www.lesswrong.com/posts/TcvcxH2Fk4n86wtoZ/anthropic-has-some-alignment-problems
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.