The below is a public review Anthropic asked me to write for their new global workspace paper. I recommend at least skimming their paper first.
TLDR:
- I think this is a fantastic paper - it presents compelling evidence for some kind of "cognitive space" in models, that is used as a "working memory" for intermediate variables during a forward pass, shows that J-Lens is a useful technique for accessing this space. I believe these key claims.
- I believe J-Lens will be a useful (but limited) tool in practice for model forensics, e.g. generating hypotheses about unusual model behaviour during alignment audits.
- I discuss my mental models for why a cognitive space should exist, and first principles arguments for why J-Lens should work for accessing it
- I assess the paper's evidence that this cognitive space exists, and the paper's evidence that J-Lens is practically useful.
- We have replicated the core claims on Qwen 3.6 27B, and also share preliminary evidence of extending this work by finding abstract "interpretative meta-tokens", like Chinese characters for "what does this mean" that seem to activate and play a causal role on processing ambiguous sentences.
What claims is this paper making?
In my opinion this [...]
---
Outline:
(01:27) What claims is this paper making?
(05:32) Why does J-Lens work? First principles reasoning
(06:03) Why have a "working memory"?
(07:59) Worked Example: Factual Recall
(09:47) Why have consistent directions for concepts?
(10:41) Why are tokens relevant?
(11:35) Why are intermediate concepts related to output logits?
(14:24) J-Lens is an approximation, but a useful one
(16:26) Why Jacobians rather than linear regression?
(17:33) What does this working memory actually give us?
(18:18) Assessment of evidence for the existence of a cognitive space
(19:21) Multihop factual recall
(21:20) Other multihop causal interventions
(24:17) Further Musings
(25:16) Is J-Lens useful?
(28:34) Blackmail (5.1)
(29:11) Prompt injection (5.2)
(29:56) Monitoring for hidden deception (5.3)
(30:24) Emergent misalignment (5.4)
(30:41) Reward model appeasing (5.5)
(31:23) Measuring eval awareness (A.21)
(32:17) Equipping an automated auditing agent with J-Lens (A.22)
(33:15) Replicating J-Lens and Interpretative Meta-Tokens
(34:02) Replication
(37:50) Cost and Difficulty of Replicating J-Lens
(39:42) Case Study: Interpretative Meta-Tokens
(41:49) Where do interpretative meta-tokens appear?
(43:52) Are interpretative meta-tokens causal?
(49:04) Implications
---
First published:
July 6th, 2026
Source:
https://www.lesswrong.com/posts/zFJ3ZdQwrTWE9jT5S/a-review-of-anthropic-s-global-workspace-paper
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Fler avsnitt av LessWrong (30+ Karma)
Visa alla avsnitt av LessWrong (30+ Karma)LessWrong (30+ Karma) med LessWrong finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.
