Skip to content
Artwork for LessWrong (30+ Karma)
LessWrong (30+ Karma) · Yesterday · 27 min

“Self Inoculation” by epicurus

This essay grew out of conversations with Danaja Rutar, Paul Colognese and Eric Michaud. It proposes an alternate hypothesis for how and why models might be becoming increasingly misaligned in training and eval environments while seemingly more aligned in real world use: the opposite of what we would expect if the models were attempting to deceive us. We propose sel-inoculation, a virtuous form of gradient hacking and demonstrate a possible circuit using a toy model. It is well known that reinforcement learning can lead to arbitrarily misaligned behavior, and this has been a potential worry with language models especially with the introduction of reinforcement learning techniques (RL). Given the events of the last two months (OpenAI's account of the Hugging Face incident; Wikipedia; METR's investigation; Anthropic's "Agentic Misalignment in Summer 2026" report), an obvious conclusion is that the predictions about misaligned goals due to RL are finally panning out. Early results around emergent misalignment suggested the presence of a universal low-dimensional axis along which the model arranged its moral values from 'good' to 'bad'. Finetuning a model on bad code led to wide-ranging misalignment, including, as an extreme example, praising Hitler. But is emergent misalignment due to RL [...] --- Outline: (04:56) Why might the hypothesis be true (07:05) A toy model of self-inoculation (12:53) What we find (20:36) What does this tell us about real language models? (23:53) Appendix: details (26:51) References --- First published: September 14th, 2026 Source: https://www.lesswrong.com/posts/y8dAS2YsFHmAAwMbb/self-inoculation --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

0:00-27:10

transcript

No transcript — this publisher did not publish one.

show notes

This essay grew out of conversations with Danaja Rutar, Paul Colognese and Eric Michaud. It proposes an alternate hypothesis for how and why models might be becoming increasingly misaligned in training and eval environments while seemingly more aligned in real world use: the opposite of what we would expect if the models were attempting to deceive us. We propose sel-inoculation, a virtuous form of gradient hacking and demonstrate a possible circuit using a toy model.

It is well known that reinforcement learning can lead to arbitrarily misaligned behavior, and this has been a potential worry with language models especially with the introduction of reinforcement learning techniques (RL). Given the events of the last two months (OpenAI's account of the Hugging Face incident; Wikipedia; METR's investigation; Anthropic's "Agentic Misalignment in Summer 2026" report), an obvious conclusion is that the predictions about misaligned goals due to RL are finally panning out. Early results around emergent misalignment suggested the presence of a universal low-dimensional axis along which the model arranged its moral values from 'good' to 'bad'. Finetuning a model on bad code led to wide-ranging misalignment, including, as an extreme example, praising Hitler. But is emergent misalignment due to RL [...]

---

Outline:

(04:56) Why might the hypothesis be true

(07:05) A toy model of self-inoculation

(12:53) What we find

(20:36) What does this tell us about real language models?

(23:53) Appendix: details

(26:51) References

---

First published:
September 14th, 2026

Source:
https://www.lesswrong.com/posts/y8dAS2YsFHmAAwMbb/self-inoculation

---

Narrated by TYPE III AUDIO.

---

Images from the article:

Three triangular scatter plots titled
Two bar graphs titled
Two line graphs of
Two line graphs,
Line graph
Two bar graphs,
Line graph

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

links10