
“WorkspaceBench: Evaluating Interpretability Methods for the Global Workspace” by camilablank, agam_bhatia, Euan Ong, Neel Nanda
Om avsnittet
TL;DR
- We introduce WorkspaceBench, a set of evaluations for how well an activation-to-text tool can read the contents of the “global workspace” of a model, i.e. the intermediate variables during a forward pass.
- The benchmark comprises 3,356 questions across 27 eval families, spanning topics in safety, logical reasoning, and multihop computation, with a subset for single-token-output tools.
- A desirable property of good interpretability techniques is minimal hallucinations, so WorkspaceBench also provides a hallucination-focused eval.
- WorkspaceBench was developed for Qwen-3.6-27B and we expect it to work on larger models, but it may need to be adapted for smaller or weaker models to ensure the models can do the tasks.
- Our goal is to create an eval that could identify a good multi-token J-lens. We open-source our benchmark here.
Introduction
Astra can do a concerning amount with no chain of thought. This is bad for CoT monitorability and makes interpretability essential to actually understanding what is going on. A key goal of interpretability is to understand intermediate variables that a model uses to compute its answers. The intermediate representations that models store in their global workspaces contain useful information that can help us decode their intentions, beliefs, algorithms [...]
---
Outline:
(00:16) TL;DR
(01:29) Introduction
(02:39) Why has no one made a WorkspaceBench before?
(04:55) Background
(08:08) WorkspaceBench
(08:12) Desiderata of a workspace reader
(10:42) Quality control
(10:46) Choosing Questions that surface intermediate variables
(12:09) Adapting WorkspaceBench to other models
(12:49) Comparing activation readers
(14:10) Grading
(14:46) Baselines
(15:55) Guarding against hallucinations
(16:52) Evaluation Sets
(17:15) Basic
(23:03) Computational
(26:13) Safety
(27:49) Association
(30:41) Anti bag of words
(31:25) Hallucination
(33:14) Logical processing
(34:08) Results
(34:11) Overall WorkspaceBench scores
(34:15) en-US-AvaMultilingualNeural__ Grouped bar graph showing pass rate across tasks for various lens methods.
(34:25) J-lens precision vs. recall for metamodels (NLAs and Oracle Lens)
(34:32) en-US-AvaMultilingualNeural__ Scatter plot showing precision against J-lens top-10 versus recall@10 across four methods.
(34:43) Discussion
(34:46) Single token readers don't surface important workspace content
(35:44) NLAs surface workspace content, but tend to hallucinate
(36:40) Acknowledgments
(36:53) Appendix
(36:56) Oracle Lens
(38:20) Agentic Evals
---
First published:
September 22nd, 2026
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Fler avsnitt
Visa alla avsnitt av LessWrong (30+ Karma)LessWrong (30+ Karma) med LessWrong finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.