″“Alignment Engineering” vs. “Misalignment Science”” by Edward James Young
transcript
show notes
There has been much discussion recently around whether a large portion of alignment research is net negative. Without endorsing or refuting them, the basic arguments here are:
- Prosaic alignment of models is becoming a bottleneck for capabilities.
- Therefore improving the prosaic alignment of models enables faster capabilities advances, which bring us closer to RSI.
- It is unlikely these prosaic alignment methods remain sufficient during the RSI loop, and so this work brings us closer to doom.
- Furthermore, dealing with these more prosaic failures reduces the likelihood of a warning shot of sufficient magnitude to cause a slowdown which would prevent RSI.
On the basis of this argument, some urge alignment researchers at AGI companies to quit outright. But quit to do what? Missing from this exchange so far has been a discussion of opportunity costs. If you aren’t going to do (technical) work on “Alignment” – either inside or outside of an AGI company – what should you work on?
In this post, I outline a contrast between “Alignment Engineering” – the dominant model for what “working on alignment” looks like (inside labs, and in the field as a whole) with “Misalignment Science”. I begin by characterising [...]
---
Outline:
(01:51) "Alignment Engineering"
(06:22) AI Safety and the ML tradition
(07:42) Implicit work trials
(09:36) "Misalignment Science"
(14:56) Conclusion
(15:39) Postscript: Iterating ourselves into oblivion
The original text contained 4 footnotes which were omitted from this narration.
---
First published:
October 5th, 2026
Source:
https://www.lesswrong.com/posts/FogmcDHA6AdMGukum/alignment-engineering-vs-misalignment-science
---
Narrated by TYPE III AUDIO.