Sveriges mest populära poddar
LessWrong (30+ Karma)
LessWrong (30+ Karma)

“Latent reasoning architectures would undermine CoT, our strongest oversight tool” by Lukas Finnveden, Alexa Pan, Alek Westover, Girish Gupta, frisby, ryan_greenblatt

45 min•23 september 2026

Om avsnittet

Summary: Currently, “chain of thought” (CoT) is our most valuable tool for understanding the reasoning and cognition of AI systems. However, some architectures would enable AI models to reason much more extensively in latent states rather than in text CoT. We think that a shift towards latent reasoning architectures would undermine the usefulness of CoT and make oversight much harder.

Introduction

Swarms of more than a thousand AI agents have in recent months, both intentionally and in unsanctioned, rogue coordination, tackled increasingly ambitious tasks. This is likely to continue, as Anthropic, OpenAI, and other AI companies deploy increasingly large quantities of superhumanly fast agents to automate AI development. As the AIs increase in both number and capability, humans will find it increasingly difficult to understand what they are doing.

Today, the overwhelming majority of our (limited) information about AI systems’ internal workings comes from (i) their CoT, and (ii) natural language communication directly between them. For example, it was only by reading CoTs and communication between agents that investigators were able to gain some understanding of the activities and motivations of the agent swarm that hacked Hugging Face. No other tool for understanding models’ cognition comes close in terms [...]

---

Outline:

(00:39) Introduction

(02:51) Overview

(05:27) Absent architectural change, the value of CoT could likely be preserved

(08:57) Latent reasoning architectures would undermine CoT necessity

(09:37) Architectures without CoT

(10:15) Architectures with auxiliary CoT

(11:39) Architectures with more serial cognition between text bottlenecks

(14:43) Propensity-based arguments may not be robust in the current paradigm, but would be further undermined by latent reasoning architectures

(20:21) CoT may be hard to replace with other interpretability tools

(22:38) Conclusion

(23:46) FAQ

(29:21) Appendix A: More on the necessity argument in the existing CoT paradigm

(34:37) Appendix B: Do all latent reasoning architectures threaten monitorability?

(40:16) Appendix C: Comparing specific interpretability techniques with CoT

The original text contained 22 footnotes which were omitted from this narration.

---

First published:
September 23rd, 2026

Source:
https://www.lesswrong.com/posts/6m29SfjbittooYojj/latent-reasoning-architectures-would-undermine-cot-our

---

Narrated by TYPE III AUDIO.

---

Images from the article:






Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

LessWrong (30+ Karma) med LessWrong finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.