“Mitigating Reward Hacking as Institutional Design” by beren
transcript
show notes
Author's Note: Cross-posted from my personal blog. The original post was on August 17th but given recent events I thought this might also be interesting to lesswrong people.
Last year I wrote a post on reward hacking as we were then beginning to see concerning signs of scaling RLVR causing models to exhibit substantial reward hacking behaviours. Unfortunately these behaviours have seemingly only grown substantially worse and more sophisticated with scale, as predicted, leading to events which cannot be described as other than egregious misalignment such as the recent OpenAI-Huggingface hacking incident. I strongly recommend everybody watch this talk presenting the details of the attack from OpenAI's perspective. It is insane.
Clearly reward hacking is now top of mind and appears to be the first potentially seriously dangerous class of misalignment that we have seen. In my original post, I described two classes of reward hacks -- 'high complexity' and 'low complexity' hacks. 'High complexity' hacks are like the early reward hacks we saw on Atari where some extremely idiosyncratic set of moves is learnt that maximizes reward in a very precise way, and can be analogized to overfitting on the reward function. 'Low' complexity hacks [...]
The original text contained 14 footnotes which were omitted from this narration.
---
First published:
September 11th, 2026
Source:
https://www.lesswrong.com/posts/tgcooi77NXMquCR5L/mitigating-reward-hacking-as-institutional-design
---
Narrated by TYPE III AUDIO.