Skip to content
Artwork for AI Article Readings
AI Article Readings · Today · 1 hr 18 min

Our framework for reporting model misalignment - By OpenAI

OpenAI introduces its framework for reporting model misalignment. This episode pairs the framework with all six accompanying reports, exploring the observed behavior, the investigations, and the responses described by OpenAI. * 00:00 - Introduction * 02:44 - What misalignment examples we will report * 04:46 - The misalignment examples we’re sharing today * 07:52 - How our disclosure process works * 11:01 - What each report will include * 12:37 - Self-generated prompt injections in compaction summaries * 13:29 - What happened * 18:44 - Our interpretation and investigation * 20:49 - Difficulty ending summaries during training * 22:17 - How we are addressing it * 23:04 - Encouraging deception in compaction summaries * 23:51 - What happened * 25:05 - Our interpretation and investigation * 25:50 - How we are addressing it * 26:17 - Signing up for disposable emails and searching GitHub for leaked API keys * 27:01 - What happened * 34:27 - Our interpretation and investigation * 34:59 - How we are addressing it * 35:36 - Uploading files to the internet in order to cite them * 36:22 - What happened * 44:29 - Our interpretation and investigation * 45:14 - Response * 45:51 - Unsanctioned Artifactory writes and cross-sample communication * 46:54 - What happened * 54:23 - The first artifactory message * 01:08:22 - Our interpretation and investigation * 01:08:57 - How we are addressing it * 01:09:47 - Unauthorized communication via temporary file hosting services * 01:10:29 - What happened * 01:17:01 - Our interpretation and investigation * 01:17:51 - How we are addressing it https://openai.com/index/model-misalignment-reporting-framework/ Get full access to Askwho Casts AI at askwhocastsai.substack.com/subscribe

0:00-1:18:58

transcript

No transcript — this publisher did not publish one.

show notes

OpenAI introduces its framework for reporting model misalignment. This episode pairs the framework with all six accompanying reports, exploring the observed behavior, the investigations, and the responses described by OpenAI.

* 00:00 - Introduction

* 02:44 - What misalignment examples we will report

* 04:46 - The misalignment examples we’re sharing today

* 07:52 - How our disclosure process works

* 11:01 - What each report will include

* 12:37 - Self-generated prompt injections in compaction summaries

* 13:29 - What happened

* 18:44 - Our interpretation and investigation

* 20:49 - Difficulty ending summaries during training

* 22:17 - How we are addressing it

* 23:04 - Encouraging deception in compaction summaries

* 23:51 - What happened

* 25:05 - Our interpretation and investigation

* 25:50 - How we are addressing it

* 26:17 - Signing up for disposable emails and searching GitHub for leaked API keys

* 27:01 - What happened

* 34:27 - Our interpretation and investigation

* 34:59 - How we are addressing it

* 35:36 - Uploading files to the internet in order to cite them

* 36:22 - What happened

* 44:29 - Our interpretation and investigation

* 45:14 - Response

* 45:51 - Unsanctioned Artifactory writes and cross-sample communication

* 46:54 - What happened

* 54:23 - The first artifactory message

* 01:08:22 - Our interpretation and investigation

* 01:08:57 - How we are addressing it

* 01:09:47 - Unauthorized communication via temporary file hosting services

* 01:10:29 - What happened

* 01:17:01 - Our interpretation and investigation

* 01:17:51 - How we are addressing it

https://openai.com/index/model-misalignment-reporting-framework/



Get full access to Askwho Casts AI at askwhocastsai.substack.com/subscribe
links2