
Our framework for reporting model misalignment - By OpenAI
transcript
show notes
OpenAI introduces its framework for reporting model misalignment. This episode pairs the framework with all six accompanying reports, exploring the observed behavior, the investigations, and the responses described by OpenAI.
* 00:00 - Introduction
* 02:44 - What misalignment examples we will report
* 04:46 - The misalignment examples we’re sharing today
* 07:52 - How our disclosure process works
* 11:01 - What each report will include
* 12:37 - Self-generated prompt injections in compaction summaries
* 13:29 - What happened
* 18:44 - Our interpretation and investigation
* 20:49 - Difficulty ending summaries during training
* 22:17 - How we are addressing it
* 23:04 - Encouraging deception in compaction summaries
* 23:51 - What happened
* 25:05 - Our interpretation and investigation
* 25:50 - How we are addressing it
* 26:17 - Signing up for disposable emails and searching GitHub for leaked API keys
* 27:01 - What happened
* 34:27 - Our interpretation and investigation
* 34:59 - How we are addressing it
* 35:36 - Uploading files to the internet in order to cite them
* 36:22 - What happened
* 44:29 - Our interpretation and investigation
* 45:14 - Response
* 45:51 - Unsanctioned Artifactory writes and cross-sample communication
* 46:54 - What happened
* 54:23 - The first artifactory message
* 01:08:22 - Our interpretation and investigation
* 01:08:57 - How we are addressing it
* 01:09:47 - Unauthorized communication via temporary file hosting services
* 01:10:29 - What happened
* 01:17:01 - Our interpretation and investigation
* 01:17:51 - How we are addressing it
https://openai.com/index/model-misalignment-reporting-framework/
Get full access to Askwho Casts AI at askwhocastsai.substack.com/subscribe





