LessWrong (30+ Karma)“Towards embedded evaluations for scheming propensities” by Dylan Bowman, Ezra Newman, Axel Højmark, Mia Hopman, Alex Meinke, Bronson Schoen, Marius Hobbhahn
TL;DR: A safety case for scheming must be able to substantiate these four key claims: scheming was not incentivized during training, strong evaluations and red teaming suggest that the model does not have a propensity to scheme, scheming was never attempted during internal deployment, and scheming reasoning would be detected (e.g., via CoT monitoring). All four of these claims should be addressed by embedded evaluators with persistent, deep access.
Introduction
A scheming AI is one that covertly and strategically pursues goals its developers didn’t intend. It might hide and protect its true capabilities and objectives from the people and processes meant to align, monitor, and control it. In 2024, Apollo Research demonstrated that frontier models were capable of scheming in controlled settings, and in 2025 we collaborated with OpenAI to determine how well their alignment techniques mitigate scheming.
Models must not be trained to scheme against humans, i.e. covertly and strategically pursue misaligned goals. Evaluations should find no propensity to scheme, models should never attempt to scheme in deployment, and their reasoning must remain legible so that scheming would be noticed if it occurred. Therefore, model developers must be able to confidently make these four claims, in order to [...]
---
Outline:
(00:43) Introduction
(04:00) Evaluating scheming risk requires persistent, deep access
(06:04) Four claims about scheming which must be verified
(06:21) Claim 1: Scheming was not incentivized during training
(09:00) Claim 1.1: Training did not directly incentivize scheming.
(09:28) Claim 1.2: Training did not incentivize goals or drives that would motivate scheming.
(11:49) Claim 1.3: The model did not attempt to subvert the training process itself.
(12:44) Claim 2: Strong evaluations and red teaming suggest that the model does not have a propensity to scheme
(14:09) Claim 2.1: Evaluations and red-teaming could not elicit scheming.
(14:47) Claim 2.2: Evaluations and red-teaming could not elicit goals or drives that would motivate scheming.
(15:37) Claim 3: Scheming was never attempted during internal deployment
(16:18) Claim 3.1: Monitors are deployed across all internal deployment contexts that present a risk of scheming.
(16:55) Claim 3.2: If scheming were to occur, monitors would detect it
(17:42) Claim 3.3: No scheming was attempted in monitored internal deployment activity.
(18:10) Claim 4: Scheming reasoning would be detected
(18:58) Claim 4.1: Reasoning relevant to scheming would be elicited and legibly verbalized in chain-of-thought.
(20:09) Embedded evaluators for alignment require persistent, deep access
(31:56) Takeaways and Next Steps
---
First published:
October 8th, 2026
Source:
https://www.lesswrong.com/posts/LZGuqoYHphet9sfZs/towards-embedded-evaluations-for-scheming-propensities
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.