Skip to content
Artwork for LessWrong (30+ Karma)
LessWrong (30+ Karma) · Friday · 14 min

“CoT controllability evals seem very under-elicited” by Jozdien

The CoTControl eval asks reasoning models to follow formatting constraints in their chain-of-thought (e.g. write in all lowercase, avoid a specific word) while solving questions. Models seem to mostly be pretty bad at this: recent models score between 0-30% with the exception of Mythos Preview. OpenAI and Anthropic have used this eval in recent system cards (GPT-5.5, Fable 5) to argue that their current frontier models are pretty bad at shaping their CoTs in ways that would reduce monitorability. I asked Claude Opus 4.6 to iterate on prompt templates for the eval, and find that using better prompts improves performance for open-source models by ~2-3 times or more (e.g. from 5.5% to 15% for GPT-OSS-120B). My guess is that more iteration would probably produce even stronger results. This suggests that the eval is heavily under-elicited, and that current CoT controllability numbers may be unrepresentative of how good models actually are at obfuscating their CoTs. As a concrete recommendation, I think researchers should optimize prompts a lot more when evaluating CoT control; I'm also excited about some kinds of training-based elicitation (such as this one). This isn't an entirely novel finding: the Fable system card mentions UK AISI raising a [...] --- Outline: (06:32) Results (06:35) Aggregate compliance (07:12) Generalization to held-out controllability tasks (09:09) Scaling patterns for few-shot prompts (10:08) Comparison with fine-tuning (10:50) Appendix A: Accuracy and reasoning length by setting (12:38) Appendix B: Per-mode results (13:13) Appendix: What the zero-shot prompts look like (14:08) Appendix C: Comparison with GEPA prompt optimization The original text contained 8 footnotes which were omitted from this narration. --- First published: September 11th, 2026 Source: https://www.lesswrong.com/posts/BbP2wCyDGdPWJ7PwP/cot-controllability-evals-seem-very-under-elicited --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

0:00-14:59

transcript

No transcript — this publisher did not publish one.

show notes

The CoTControl eval asks reasoning models to follow formatting constraints in their chain-of-thought (e.g. write in all lowercase, avoid a specific word) while solving questions. Models seem to mostly be pretty bad at this: recent models score between 0-30% with the exception of Mythos Preview. OpenAI and Anthropic have used this eval in recent system cards (GPT-5.5, Fable 5) to argue that their current frontier models are pretty bad at shaping their CoTs in ways that would reduce monitorability.

I asked Claude Opus 4.6 to iterate on prompt templates for the eval, and find that using better prompts improves performance for open-source models by ~2-3 times or more (e.g. from 5.5% to 15% for GPT-OSS-120B). My guess is that more iteration would probably produce even stronger results.

This suggests that the eval is heavily under-elicited, and that current CoT controllability numbers may be unrepresentative of how good models actually are at obfuscating their CoTs. As a concrete recommendation, I think researchers should optimize prompts a lot more when evaluating CoT control; I'm also excited about some kinds of training-based elicitation (such as this one).

This isn't an entirely novel finding: the Fable system card mentions UK AISI raising a [...]

---

Outline:

(06:32) Results

(06:35) Aggregate compliance

(07:12) Generalization to held-out controllability tasks

(09:09) Scaling patterns for few-shot prompts

(10:08) Comparison with fine-tuning

(10:50) Appendix A: Accuracy and reasoning length by setting

(12:38) Appendix B: Per-mode results

(13:13) Appendix: What the zero-shot prompts look like

(14:08) Appendix C: Comparison with GEPA prompt optimization

The original text contained 8 footnotes which were omitted from this narration.

---

First published:
September 11th, 2026

Source:
https://www.lesswrong.com/posts/BbP2wCyDGdPWJ7PwP/cot-controllability-evals-seem-very-under-elicited

---

Narrated by TYPE III AUDIO.

---

Images from the article:

Bar graph
Bar graph
Bar graph
Bar graph
Bar graph
Bar graph titled
Bar graph titled
Bar graph,
Line graphs titled
Bar graph,
Grouped bar graph

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

links14