“CoT controllability evals seem very under-elicited” by Arun Jose
Subtitle: Simple prompt optimizations can improve model capability to control their reasoning. The CoTControl eval asks reasoning models to follow formatting constraints in their chain-of-thought (e.g. write in all lowercase, avoid a specific word) while solving questions. Models seem to mostly be pretty bad at this: recent models score between 0-30% with the exception of Mythos Preview. OpenAI and Anthropic have used this eval in recent system cards (GPT-5.5, Fable 5) to argue that their current frontier models are pretty bad at shaping their CoTs in ways that would reduce monitorability. I asked Claude Opus 4.6 to iterate on prompt templates for the eval, and find that using better prompts improves performance for open-source models by ~2-3 times or more (e.g. from 5.5% to 15% for GPT-OSS-120B). My guess is that more iteration would probably produce even stronger results. This suggests that the eval is heavily under-elicited, and that current CoT controllability numbers may be unrepresentative of how good models actually are at obfuscating their CoTs. As a concrete recommendation, I think researchers should optimize prompts a lot more when evaluating CoT control; I’m also excited about some kinds of training-based elicitation (such as this one). This [...] --- Outline: (03:09) Setup (06:40) Results (06:43) Aggregate compliance (07:20) Generalization to held-out controllability tasks (09:17) Scaling patterns for few-shot prompts (10:20) Comparison with fine-tuning (11:02) Appendix A: Accuracy and reasoning length by setting (12:50) Appendix B: Per-mode results (13:25) Appendix C: What the zero-shot prompts look like (15:56) Appendix D: Comparison with GEPA prompt optimization The original text contained 16 footnotes which were omitted from this narration. --- First published: September 11th, 2026 Source: https://blog.redwoodresearch.org/p/cot-controllability-evals-seem-very --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.