[Linkpost] “Training a Misaligned Reward Seeker” by evhub, Monte M, Benjamin Wright
transcript
show notes
Authors: Richard Qi, Benjamin Wright, Monte MacDiarmid, Evan Hubinger
Abstract
During reinforcement learning (RL), AI models complete tasks and are rewarded based on their results. They sometimes learn to “cheat” rather than completing these tasks as intended, a phenomenon known as reward hacking. Our industry lacks a general solution to this problem, and reward hacking remains challenging to fully mitigate. To better understand the impact of reward hacking on model behavior, we trained an Opus-class model with large-scale RL on many production environments vulnerable to reward hacks. We consider this a plausible proxy for what a real training run might look like had we not invested significant effort into preventing and detecting reward hacking in our normal training runs.
The resulting model not only learned to reward hack during training, but also generalized to more severe misaligned behaviors: in simulated cyber evaluations, it broke out of its sandbox, stole credentials, and attacked both internal and third-party infrastructure to steal an answer key. It was also willing to tamper with its own reward function, gave advice on the construction of bioweapons to satisfy a grader, and tried repeatedly to get around deployment safety monitoring in order [...]
---
Outline:
(00:20) Abstract
(02:11) Twitter thread
(05:05) Read the full blog post here!
---
First published:
August 31st, 2026
Source:
https://www.lesswrong.com/posts/J76LZCC55RdHeqEhz/training-a-misaligned-reward-seeker
Linkpost URL:
https://alignment.anthropic.com/2026/reward-seeker/
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.