transcript
show notes
This essay grew out of conversations with Danaja Rutar, Paul Colognese and Eric Michaud. It proposes an alternate hypothesis for how and why models might be becoming increasingly misaligned in training and eval environments while seemingly more aligned in real world use: the opposite of what we would expect if the models were attempting to deceive us. We propose sel-inoculation, a virtuous form of gradient hacking and demonstrate a possible circuit using a toy model.
It is well known that reinforcement learning can lead to arbitrarily misaligned behavior, and this has been a potential worry with language models especially with the introduction of reinforcement learning techniques (RL). Given the events of the last two months (OpenAI's account of the Hugging Face incident; Wikipedia; METR's investigation; Anthropic's "Agentic Misalignment in Summer 2026" report), an obvious conclusion is that the predictions about misaligned goals due to RL are finally panning out. Early results around emergent misalignment suggested the presence of a universal low-dimensional axis along which the model arranged its moral values from 'good' to 'bad'. Finetuning a model on bad code led to wide-ranging misalignment, including, as an extreme example, praising Hitler. But is emergent misalignment due to RL [...]
---
Outline:
(04:56) Why might the hypothesis be true
(07:05) A toy model of self-inoculation
(12:53) What we find
(20:36) What does this tell us about real language models?
(23:53) Appendix: details
(26:51) References
---
First published:
September 14th, 2026
Source:
https://www.lesswrong.com/posts/y8dAS2YsFHmAAwMbb/self-inoculation
---
Narrated by TYPE III AUDIO.
---
Images from the article:







Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.