Sveriges mest populära poddar
Redwood Research Blog
Redwood Research Blog

“Latent reasoning architectures would undermine CoT, our strongest oversight tool” by Lukas Finnveden, Alexa Pan, Alek Westover, Girish Gupta, Nathan Sheffield, Ryan Greenblatt

54 min•23 september 2026

Om avsnittet

Subtitle: We should have a strong presumption that latent reasoning architectures would make oversight far more difficult.

Summary: Currently, “chain of thought” (CoT) is our most valuable tool for understanding the reasoning and cognition of AI systems. However, some architectures would enable AI models to reason much more extensively in latent states rather than in text CoT. We think that a shift towards latent reasoning architectures would undermine the usefulness of CoT and make oversight much harder.

Introduction

Swarms of more than a thousand AI agents have in recent months, both intentionally and in unsanctioned, rogue coordination, tackled increasingly ambitious tasks. This is likely to continue, as Anthropic, OpenAI, and other AI companies deploy increasingly large quantities of superhumanly fast agents to automate AI development. As the AIs increase in both number and capability, humans will find it increasingly difficult to understand what they are doing.

Today, the overwhelming majority of our (limited) information about AI systems’ internal workings comes from (i) their CoT, and (ii) natural language communication directly between them. For example, it was only by reading CoTs and communication between agents that investigators were able to gain some understanding of the activities and motivations of [...]

---

Outline:

(00:48) Introduction

(03:00) Overview

(05:36) Absent architectural change, the value of CoT could likely be preserved

(09:05) Latent reasoning architectures would undermine CoT necessity

(09:45) Architectures without CoT

(10:23) Architectures with auxiliary CoT

(11:47) Architectures with more serial cognition between text bottlenecks

(14:50) Propensity-based arguments may not be robust in the current paradigm, but would be further undermined by latent reasoning architectures

(20:26) CoT may be hard to replace with other interpretability tools

(22:44) Conclusion

(23:51) FAQ

(29:26) Appendix A: More on the necessity argument in the existing CoT paradigm

(34:43) Appendix B: Do all latent reasoning architectures threaten monitorability?

(40:21) Appendix C: Comparing specific interpretability techniques with CoT

The original text contained 44 footnotes which were omitted from this narration.

---

First published:
September 23rd, 2026

Source:
https://blog.redwoodresearch.org/p/latent-reasoning-architectures-would

---

Narrated by TYPE III AUDIO.

---

Images from the article:






Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Redwood Research Blog med Redwood Research finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.