“When is Unlimited Optimization Catastrophic?” by Winter Cross
transcript
show notes
This post discusses research I've completed along with my colleagues Leo Cymbalista, Alfred Harwood, and Jose Faustino at Dovetail Research. Most of the ideas in this post are expanded upon in our paper which can be found on arXiv. This work was funded by the Advanced Research + Invention Agency (ARIA) through project code MSAI-SE01-P005.
A common justification for the danger of AI comes from the idea that human value is fragile. That is, if we modify our values and heavily optimize the world for the modification, we are likely to end up in a valueless world. In the LessWrong post Value is Fragile which canonicalizes this idea, Eliezer Yudkowsky gives several examples where "forgetting" to specify a dimension of human value such as consciousness or boredom to a powerful AI can intuitively result in an undesirable outcome that is endlessly repetitive or meaningless respectively. While his examples in the post all take this form, he argues more generally that any future not shaped with reliable inheritance from human values will contain almost nothing of worth. This idea is especially concerning in the midst of current-day AIs aligned through one-time techniques such as RLHF before being deployed [...]
---
Outline:
(02:10) A Model of Alignment
(05:32) Alignment Tests
(05:56) Finite Framework
(07:02) Continuous Framework
(08:08) Attributes Framework
(11:19) Results
(11:22) Finite Framework
(12:34) Example
(13:58) Continuous Framework
(15:36) Example
(17:00) Attributes Framework
(19:31) Example
(21:01) Discussion
(21:53) Future Work
The original text contained 1 footnote which was omitted from this narration.
---
First published:
August 21st, 2026
Source:
https://www.lesswrong.com/posts/4JCne6evQjtjxXKED/when-is-unlimited-optimization-catastrophic
---
Narrated by TYPE III AUDIO.
---
Images from the article:

![An image showing the graph of two value functions over the set of world states [0,1]. Left: A graph displaying the true/human value function graphed as a blue line. Right: A graph displaying the proxy value function graphed as a black line. This proxy is catastrophic since its unique highest-valued state is exactly the lowest-valued state of the true value function.](https://res.cloudinary.com/lesswrong-2-0/image/upload/v1787331124/lexical_client_uploads/sitiarsezxh4ffrtmvi9.png)


Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
- https://www.lesswrong.com/posts/4JCne6evQjtjxXKED/when-is-unlimited-optimization-catastrophiclesswrong.com
- TYPE III AUDIOtype3.audio
- A diagram showing a railroad junction. Two trains are traveling to the right and the agent can control two switches to alter their courses. There are four states the agent can leave the railroad in: both switches set to straight (aa), train 2 made to turn only (ab), train 1 made to turn only (ba), and both trains made to turn (bb).res.cloudinary.com
- An image showing the graph of two value functions over the set of world states [0,1]. Left: A graph displaying the true/human value function graphed as a blue line. Right: A graph displaying the proxy value function graphed as a black line. This proxy is catastrophic since its unique highest-valued state is exactly the lowest-valued state of the true value function.res.cloudinary.com