Skip to content
Artwork for LessWrong (30+ Karma)
LessWrong (30+ Karma) · Sunday · 56 min

“Mitigating Reward Hacking as Institutional Design” by beren

Author's Note: Cross-posted from my personal blog. The original post was on August 17th but given recent events I thought this might also be interesting to lesswrong people. Last year I wrote a post on reward hacking as we were then beginning to see concerning signs of scaling RLVR causing models to exhibit substantial reward hacking behaviours. Unfortunately these behaviours have seemingly only grown substantially worse and more sophisticated with scale, as predicted, leading to events which cannot be described as other than egregious misalignment such as the recent OpenAI-Huggingface hacking incident. I strongly recommend everybody watch this talk presenting the details of the attack from OpenAI's perspective. It is insane. Clearly reward hacking is now top of mind and appears to be the first potentially seriously dangerous class of misalignment that we have seen. In my original post, I described two classes of reward hacks -- 'high complexity' and 'low complexity' hacks. 'High complexity' hacks are like the early reward hacks we saw on Atari where some extremely idiosyncratic set of moves is learnt that maximizes reward in a very precise way, and can be analogized to overfitting on the reward function. 'Low' complexity hacks [...] The original text contained 14 footnotes which were omitted from this narration. --- First published: September 11th, 2026 Source: https://www.lesswrong.com/posts/tgcooi77NXMquCR5L/mitigating-reward-hacking-as-institutional-design --- Narrated by TYPE III AUDIO.

0:00-56:47

transcript

No transcript — this publisher did not publish one.

show notes

Author's Note: Cross-posted from my personal blog. The original post was on August 17th but given recent events I thought this might also be interesting to lesswrong people.


Last year I wrote a post on reward hacking as we were then beginning to see concerning signs of scaling RLVR causing models to exhibit substantial reward hacking behaviours. Unfortunately these behaviours have seemingly only grown substantially worse and more sophisticated with scale, as predicted, leading to events which cannot be described as other than egregious misalignment such as the recent OpenAI-Huggingface hacking incident. I strongly recommend everybody watch this talk presenting the details of the attack from OpenAI's perspective. It is insane.

Clearly reward hacking is now top of mind and appears to be the first potentially seriously dangerous class of misalignment that we have seen. In my original post, I described two classes of reward hacks -- 'high complexity' and 'low complexity' hacks. 'High complexity' hacks are like the early reward hacks we saw on Atari where some extremely idiosyncratic set of moves is learnt that maximizes reward in a very precise way, and can be analogized to overfitting on the reward function. 'Low' complexity hacks [...]

The original text contained 14 footnotes which were omitted from this narration.

---

First published:
September 11th, 2026

Source:
https://www.lesswrong.com/posts/tgcooi77NXMquCR5L/mitigating-reward-hacking-as-institutional-design

---

Narrated by TYPE III AUDIO.

links2