
“CoT controllability evals seem very under-elicited” by Arun Jose
Om avsnittet
Subtitle: Simple prompt optimizations can improve model capability to control their reasoning.
The CoTControl eval asks reasoning models to follow formatting constraints in their chain-of-thought (e.g. write in all lowercase, avoid a specific word) while solving questions. Models seem to mostly be pretty bad at this: recent models score between 0-30% with the exception of Mythos Preview. OpenAI and Anthropic have used this eval in recent system cards (GPT-5.5, Fable 5) to argue that their current frontier models are pretty bad at shaping their CoTs in ways that would reduce monitorability.
I asked Claude Opus 4.6 to iterate on prompt templates for the eval, and find that using better prompts improves performance for open-source models by ~2-3 times or more (e.g. from 5.5% to 15% for GPT-OSS-120B). My guess is that more iteration would probably produce even stronger results.
This suggests that the eval is heavily under-elicited, and that current CoT controllability numbers may be unrepresentative of how good models actually are at obfuscating their CoTs. As a concrete recommendation, I think researchers should optimize prompts a lot more when evaluating CoT control; I’m also excited about some kinds of training-based elicitation (such as this one).
This [...]
---
Outline:
(03:09) Setup
(06:40) Results
(06:43) Aggregate compliance
(07:20) Generalization to held-out controllability tasks
(09:17) Scaling patterns for few-shot prompts
(10:20) Comparison with fine-tuning
(11:02) Appendix A: Accuracy and reasoning length by setting
(12:50) Appendix B: Per-mode results
(13:25) Appendix C: What the zero-shot prompts look like
(15:56) Appendix D: Comparison with GEPA prompt optimization
The original text contained 16 footnotes which were omitted from this narration.
---
First published:
September 11th, 2026
Source:
https://blog.redwoodresearch.org/p/cot-controllability-evals-seem-very
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Fler avsnitt
Visa alla avsnitt av Redwood Research BlogRedwood Research Blog med Redwood Research finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.