
“Latent reasoning architectures would undermine CoT, our strongest oversight tool” by Lukas Finnveden, Alexa Pan, Alek Westover, Girish Gupta, Nathan Sheffield, Ryan Greenblatt
Om avsnittet
Subtitle: We should have a strong presumption that latent reasoning architectures would make oversight far more difficult.
Summary: Currently, “chain of thought” (CoT) is our most valuable tool for understanding the reasoning and cognition of AI systems. However, some architectures would enable AI models to reason much more extensively in latent states rather than in text CoT. We think that a shift towards latent reasoning architectures would undermine the usefulness of CoT and make oversight much harder.
Introduction
Swarms of more than a thousand AI agents have in recent months, both intentionally and in unsanctioned, rogue coordination, tackled increasingly ambitious tasks. This is likely to continue, as Anthropic, OpenAI, and other AI companies deploy increasingly large quantities of superhumanly fast agents to automate AI development. As the AIs increase in both number and capability, humans will find it increasingly difficult to understand what they are doing.
Today, the overwhelming majority of our (limited) information about AI systems’ internal workings comes from (i) their CoT, and (ii) natural language communication directly between them. For example, it was only by reading CoTs and communication between agents that investigators were able to gain some understanding of the activities and motivations of [...]
---
Outline:
(00:48) Introduction
(03:00) Overview
(05:36) Absent architectural change, the value of CoT could likely be preserved
(09:05) Latent reasoning architectures would undermine CoT necessity
(09:45) Architectures without CoT
(10:23) Architectures with auxiliary CoT
(11:47) Architectures with more serial cognition between text bottlenecks
(14:50) Propensity-based arguments may not be robust in the current paradigm, but would be further undermined by latent reasoning architectures
(20:26) CoT may be hard to replace with other interpretability tools
(22:44) Conclusion
(23:51) FAQ
(29:26) Appendix A: More on the necessity argument in the existing CoT paradigm
(34:43) Appendix B: Do all latent reasoning architectures threaten monitorability?
(40:21) Appendix C: Comparing specific interpretability techniques with CoT
The original text contained 44 footnotes which were omitted from this narration.
---
First published:
September 23rd, 2026
Source:
https://blog.redwoodresearch.org/p/latent-reasoning-architectures-would
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Fler avsnitt
Visa alla avsnitt av Redwood Research BlogRedwood Research Blog med Redwood Research finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.