Sveriges mest populära poddar
LessWrong (30+ Karma)
LessWrong (30+ Karma)

“WorkspaceBench: Evaluating Interpretability Methods for the Global Workspace” by camilablank, agam_bhatia, Euan Ong, Neel Nanda

39 min•24 september 2026

Om avsnittet

TL;DR

  • We introduce WorkspaceBench, a set of evaluations for how well an activation-to-text tool can read the contents of the “global workspace” of a model, i.e. the intermediate variables during a forward pass.
  • The benchmark comprises 3,356 questions across 27 eval families, spanning topics in safety, logical reasoning, and multihop computation, with a subset for single-token-output tools.
  • A desirable property of good interpretability techniques is minimal hallucinations, so WorkspaceBench also provides a hallucination-focused eval.
  • WorkspaceBench was developed for Qwen-3.6-27B and we expect it to work on larger models, but it may need to be adapted for smaller or weaker models to ensure the models can do the tasks.
  • Our goal is to create an eval that could identify a good multi-token J-lens. We open-source our benchmark here.



Introduction

Astra can do a concerning amount with no chain of thought. This is bad for CoT monitorability and makes interpretability essential to actually understanding what is going on. A key goal of interpretability is to understand intermediate variables that a model uses to compute its answers. The intermediate representations that models store in their global workspaces contain useful information that can help us decode their intentions, beliefs, algorithms [...]

---

Outline:

(00:16) TL;DR

(01:29) Introduction

(02:39) Why has no one made a WorkspaceBench before?

(04:55) Background

(08:08) WorkspaceBench

(08:12) Desiderata of a workspace reader

(10:42) Quality control

(10:46) Choosing Questions that surface intermediate variables

(12:09) Adapting WorkspaceBench to other models

(12:49) Comparing activation readers

(14:10) Grading

(14:46) Baselines

(15:55) Guarding against hallucinations

(16:52) Evaluation Sets

(17:15) Basic

(23:03) Computational

(26:13) Safety

(27:49) Association

(30:41) Anti bag of words

(31:25) Hallucination

(33:14) Logical processing

(34:08) Results

(34:11) Overall WorkspaceBench scores

(34:15) en-US-AvaMultilingualNeural__ Grouped bar graph showing pass rate across tasks for various lens methods.

(34:25) J-lens precision vs. recall for metamodels (NLAs and Oracle Lens)

(34:32) en-US-AvaMultilingualNeural__ Scatter plot showing precision against J-lens top-10 versus recall@10 across four methods.

(34:43) Discussion

(34:46) Single token readers don't surface important workspace content

(35:44) NLAs surface workspace content, but tend to hallucinate

(36:40) Acknowledgments

(36:53) Appendix

(36:56) Oracle Lens

(38:20) Agentic Evals

---

First published:
September 22nd, 2026

Source:
https://www.lesswrong.com/posts/Zeg2JztbdhguL48uH/workspacebench-evaluating-interpretability-methods-for-the

---

Narrated by TYPE III AUDIO.

---

Images from the article:




Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

LessWrong (30+ Karma) med LessWrong finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.