Sveriges mest populära poddar
LessWrong (30+ Karma)

“A Review of Anthropic’s Global Workspace Paper” by Neel Nanda

51 min6 juli 2026

The below is a public review Anthropic asked me to write for their new global workspace paper. I recommend at least skimming their paper first.

TLDR:

  • I think this is a fantastic paper - it presents compelling evidence for some kind of "cognitive space" in models, that is used as a "working memory" for intermediate variables during a forward pass, shows that J-Lens is a useful technique for accessing this space. I believe these key claims.
  • I believe J-Lens will be a useful (but limited) tool in practice for model forensics, e.g. generating hypotheses about unusual model behaviour during alignment audits.
  • I discuss my mental models for why a cognitive space should exist, and first principles arguments for why J-Lens should work for accessing it
  • I assess the paper's evidence that this cognitive space exists, and the paper's evidence that J-Lens is practically useful.
  • We have replicated the core claims on Qwen 3.6 27B, and also share preliminary evidence of extending this work by finding abstract "interpretative meta-tokens", like Chinese characters for "what does this mean" that seem to activate and play a causal role on processing ambiguous sentences.

What claims is this paper making?

In my opinion this [...]

---

Outline:

(01:27) What claims is this paper making?

(05:32) Why does J-Lens work? First principles reasoning

(06:03) Why have a "working memory"?

(07:59) Worked Example: Factual Recall

(09:47) Why have consistent directions for concepts?

(10:41) Why are tokens relevant?

(11:35) Why are intermediate concepts related to output logits?

(14:24) J-Lens is an approximation, but a useful one

(16:26) Why Jacobians rather than linear regression?

(17:33) What does this working memory actually give us?

(18:18) Assessment of evidence for the existence of a cognitive space

(19:21) Multihop factual recall

(21:20) Other multihop causal interventions

(24:17) Further Musings

(25:16) Is J-Lens useful?

(28:34) Blackmail (5.1)

(29:11) Prompt injection (5.2)

(29:56) Monitoring for hidden deception (5.3)

(30:24) Emergent misalignment (5.4)

(30:41) Reward model appeasing (5.5)

(31:23) Measuring eval awareness (A.21)

(32:17) Equipping an automated auditing agent with J-Lens (A.22)

(33:15) Replicating J-Lens and Interpretative Meta-Tokens

(34:02) Replication

(37:50) Cost and Difficulty of Replicating J-Lens

(39:42) Case Study: Interpretative Meta-Tokens

(41:49) Where do interpretative meta-tokens appear?

(43:52) Are interpretative meta-tokens causal?

(49:04) Implications

---

First published:
July 6th, 2026

Source:
https://www.lesswrong.com/posts/zFJ3ZdQwrTWE9jT5S/a-review-of-anthropic-s-global-workspace-paper

---

Narrated by TYPE III AUDIO.

---

Images from the article:

















Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Fler avsnitt av LessWrong (30+ Karma)

Visa alla avsnitt av LessWrong (30+ Karma)

LessWrong (30+ Karma) med LessWrong finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.