Skip to content
Artwork for LessWrong (30+ Karma)
LessWrong (30+ Karma) · August 22 · 23 min

“When is Unlimited Optimization Catastrophic?” by Winter Cross

This post discusses research I've completed along with my colleagues Leo Cymbalista, Alfred Harwood, and Jose Faustino at Dovetail Research. Most of the ideas in this post are expanded upon in our paper which can be found on arXiv. This work was funded by the Advanced Research + Invention Agency (ARIA) through project code MSAI-SE01-P005. A common justification for the danger of AI comes from the idea that human value is fragile. That is, if we modify our values and heavily optimize the world for the modification, we are likely to end up in a valueless world. In the LessWrong post Value is Fragile which canonicalizes this idea, Eliezer Yudkowsky gives several examples where "forgetting" to specify a dimension of human value such as consciousness or boredom to a powerful AI can intuitively result in an undesirable outcome that is endlessly repetitive or meaningless respectively. While his examples in the post all take this form, he argues more generally that any future not shaped with reliable inheritance from human values will contain almost nothing of worth. This idea is especially concerning in the midst of current-day AIs aligned through one-time techniques such as RLHF before being deployed [...] --- Outline: (02:10) A Model of Alignment (05:32) Alignment Tests (05:56) Finite Framework (07:02) Continuous Framework (08:08) Attributes Framework (11:19) Results (11:22) Finite Framework (12:34) Example (13:58) Continuous Framework (15:36) Example (17:00) Attributes Framework (19:31) Example (21:01) Discussion (21:53) Future Work The original text contained 1 footnote which was omitted from this narration. --- First published: August 21st, 2026 Source: https://www.lesswrong.com/posts/4JCne6evQjtjxXKED/when-is-unlimited-optimization-catastrophic --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

0:00-23:58

transcript

No transcript — this publisher did not publish one.

show notes

This post discusses research I've completed along with my colleagues Leo Cymbalista, Alfred Harwood, and Jose Faustino at Dovetail Research. Most of the ideas in this post are expanded upon in our paper which can be found on arXiv. This work was funded by the Advanced Research + Invention Agency (ARIA) through project code MSAI-SE01-P005.


A common justification for the danger of AI comes from the idea that human value is fragile. That is, if we modify our values and heavily optimize the world for the modification, we are likely to end up in a valueless world. In the LessWrong post Value is Fragile which canonicalizes this idea, Eliezer Yudkowsky gives several examples where "forgetting" to specify a dimension of human value such as consciousness or boredom to a powerful AI can intuitively result in an undesirable outcome that is endlessly repetitive or meaningless respectively. While his examples in the post all take this form, he argues more generally that any future not shaped with reliable inheritance from human values will contain almost nothing of worth. This idea is especially concerning in the midst of current-day AIs aligned through one-time techniques such as RLHF before being deployed [...]

---

Outline:

(02:10) A Model of Alignment

(05:32) Alignment Tests

(05:56) Finite Framework

(07:02) Continuous Framework

(08:08) Attributes Framework

(11:19) Results

(11:22) Finite Framework

(12:34) Example

(13:58) Continuous Framework

(15:36) Example

(17:00) Attributes Framework

(19:31) Example

(21:01) Discussion

(21:53) Future Work

The original text contained 1 footnote which was omitted from this narration.

---

First published:
August 21st, 2026

Source:
https://www.lesswrong.com/posts/4JCne6evQjtjxXKED/when-is-unlimited-optimization-catastrophic

---

Narrated by TYPE III AUDIO.

---

Images from the article:

A diagram showing a railroad junction. Two trains are traveling to the right and the agent can control two switches to alter their courses. There are four states the agent can leave the railroad in: both switches set to straight (aa), train 2 made to turn only (ab), train 1 made to turn only (ba), and both trains made to turn (bb).
An image showing the graph of two value functions over the set of world states [0,1]. Left: A graph displaying the true/human value function graphed as a blue line. Right: A graph displaying the proxy value function graphed as a black line. This proxy is catastrophic since its unique highest-valued state is exactly the lowest-valued state of the true value function.
A grid showing the effect of a Boltzmann optimizer on a distribution over the states. In the background of each graph is the true value function shown in blue. Graphs in the top row show the distribution (scaled to fit the box) in green when the target is the true value function while the bottom row graphs show the distribution in orange when the target is the proxy value function. Graphs in the same row show the distribution at the same optimizing power. When there is lots of optimizing power, the distributions become extremely concentrated.
A diagram showing a feasible region for two attributes. The feasible region is shaded red showing the physically possible combinations of meals prepared and customer satisfaction. Shown as blue lines are the level curves of a strictly increasing function. Despite increasing with both attributes, 's highest-value point within the feasible region is on the bottom-right where customer satisfaction is completely sacrificed to optimize meals served.

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

links7