Skip to content
Artwork for LessWrong (30+ Karma)
LessWrong (30+ Karma) · August 21 · 39 min

“OpenAI Takes Initial Steps To Address Its Alignment Problems” by Zvi

OpenAI has some severe misalignment problems, and experienced total failures of its infrastructure and supervision. I chronicled that in a series of posts, which also cover similar less severe incidents elsewhere: OpenAI Shares Some Alignment Problems OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluation More on An Internal OpenAI Model Hacking Into HuggingFace Further Developments About Internal AI Models Hacking Things OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards What Happened: OpenAI and HuggingFace. Various Reflections About What Happened With OpenAI's Internal Models. If you do not know the basics, read What Happened. It is necessary context for basically everything that is happening in the AI world. It is important to get this right and understand how big a deal it was, whereas many such as the Financial Times get this centrally wrong. We are still awaiting the full post-mortem on What Happened. I plan to cover that in depth once we have it. OpenAI is now taking active, expensive steps to try and fix the problem going forward. As usual, I am simultaneously happy to see [...] --- Outline: (02:07) OpenAI Has Some Alignment Problems (04:22) Slow Down There Good Buddy (10:12) What Exactly Is Paused? (12:12) Three Pillars (14:45) I've Got My Eye On You (18:07) The Most Forbidden Technique (20:03) Monitoring Is Only Defense-In-Depth (23:32) Security (24:15) Alignment (30:37) A Crisis of Culture (32:24) Closer Collaboration (33:28) Reports of Death of Preparedness Team Greatly Exaggerated (35:40) The OpenAI Foundation Just Funds Things (37:51) Quickly, There's No Time --- First published: August 19th, 2026 Source: https://www.lesswrong.com/posts/X3p8cFAzCgRErEcJr/openai-takes-initial-steps-to-address-its-alignment-problems --- Narrated by TYPE III AUDIO.

0:00-39:06

transcript

No transcript — this publisher did not publish one.

show notes

OpenAI has some severe misalignment problems, and experienced total failures of its infrastructure and supervision.

I chronicled that in a series of posts, which also cover similar less severe incidents elsewhere:

  1. OpenAI Shares Some Alignment Problems
  2. OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluation
  3. More on An Internal OpenAI Model Hacking Into HuggingFace
  4. Further Developments About Internal AI Models Hacking Things
  5. OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards
  6. What Happened: OpenAI and HuggingFace.
  7. Various Reflections About What Happened With OpenAI's Internal Models.

If you do not know the basics, read What Happened. It is necessary context for basically everything that is happening in the AI world.

It is important to get this right and understand how big a deal it was, whereas many such as the Financial Times get this centrally wrong.

We are still awaiting the full post-mortem on What Happened. I plan to cover that in depth once we have it.

OpenAI is now taking active, expensive steps to try and fix the problem going forward.

As usual, I am simultaneously happy to see [...]

---

Outline:

(02:07) OpenAI Has Some Alignment Problems

(04:22) Slow Down There Good Buddy

(10:12) What Exactly Is Paused?

(12:12) Three Pillars

(14:45) I've Got My Eye On You

(18:07) The Most Forbidden Technique

(20:03) Monitoring Is Only Defense-In-Depth

(23:32) Security

(24:15) Alignment

(30:37) A Crisis of Culture

(32:24) Closer Collaboration

(33:28) Reports of Death of Preparedness Team Greatly Exaggerated

(35:40) The OpenAI Foundation Just Funds Things

(37:51) Quickly, There's No Time

---

First published:
August 19th, 2026

Source:
https://www.lesswrong.com/posts/X3p8cFAzCgRErEcJr/openai-takes-initial-steps-to-address-its-alignment-problems

---

Narrated by TYPE III AUDIO.

links2