“CoT controllability evals seem very under-elicited” by Jozdien
transcript
show notes
The CoTControl eval asks reasoning models to follow formatting constraints in their chain-of-thought (e.g. write in all lowercase, avoid a specific word) while solving questions. Models seem to mostly be pretty bad at this: recent models score between 0-30% with the exception of Mythos Preview. OpenAI and Anthropic have used this eval in recent system cards (GPT-5.5, Fable 5) to argue that their current frontier models are pretty bad at shaping their CoTs in ways that would reduce monitorability.
I asked Claude Opus 4.6 to iterate on prompt templates for the eval, and find that using better prompts improves performance for open-source models by ~2-3 times or more (e.g. from 5.5% to 15% for GPT-OSS-120B). My guess is that more iteration would probably produce even stronger results.
This suggests that the eval is heavily under-elicited, and that current CoT controllability numbers may be unrepresentative of how good models actually are at obfuscating their CoTs. As a concrete recommendation, I think researchers should optimize prompts a lot more when evaluating CoT control; I'm also excited about some kinds of training-based elicitation (such as this one).
This isn't an entirely novel finding: the Fable system card mentions UK AISI raising a [...]
---
Outline:
(06:32) Results
(06:35) Aggregate compliance
(07:12) Generalization to held-out controllability tasks
(09:09) Scaling patterns for few-shot prompts
(10:08) Comparison with fine-tuning
(10:50) Appendix A: Accuracy and reasoning length by setting
(12:38) Appendix B: Per-mode results
(13:13) Appendix: What the zero-shot prompts look like
(14:08) Appendix C: Comparison with GEPA prompt optimization
The original text contained 8 footnotes which were omitted from this narration.
---
First published:
September 11th, 2026
Source:
https://www.lesswrong.com/posts/BbP2wCyDGdPWJ7PwP/cot-controllability-evals-seem-very-under-elicited
---
Narrated by TYPE III AUDIO.
---
Images from the article:











Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.