Sveriges mest populära poddar
LessWrong (30+ Karma)
LessWrong (30+ Karma)

“Exploration Hacking in AI Debate: Initial Empirics and Generalisation Splitting” by Jason R Brown, Nathalie Kirch, Joschka Braun, hyannakoudakis, David Lindner

28 min•9 september 2026

Om avsnittet

This is the first of two posts resulting from a recent Astra/MATS research project investigating exploration hacking in AI debate. They are designed to be standalone, but we encourage interested readers to read both. This post focuses on our empirical results, the second focuses on a new conceptual framework.

Authors

Jason Brown*, Nathalie Kirch*, Joschka Braun, Helen Yannakoudakis, Roland S. Zimmermann, David Lindner

*Equal contribution.

TL;DR

We set out to build model organisms of exploration hacking (EH) in the setting of AI debate. We did this by trying to create models that persistently sandbagged on certain question topics, but not on others. We ran two experiments, one to try and isolate the effects of the judge, and the other to better approximate the full dynamics of RL training on AI debates.

Overall our results indicate that EH could be a significant issue within AI debate, with the experiments respectively showing that weaker judges and longer debates slow down improvements in performance. Interestingly, the dominant mechanism behind the second result appears to be one we have not seen described before. Once the debaters were instructed to sandbag on a targeted topic, training improvements stopped transferring between targeted and non-targeted topics [...]

---

Outline:

(00:35) Authors

(00:49) TL;DR

(02:02) Introduction

(04:47) Experiment 1: Isolated Imperfect Judge

(05:17) Setup

(13:05) Experiment 2: Self-play RL on AI Debate

(13:26) Setup

(16:09) Results

(21:50) Generalisation Splitting

(25:28) Discussion

(26:47) Acknowledgements

The original text contained 6 footnotes which were omitted from this narration.

---

First published:
September 8th, 2026

Source:
https://www.lesswrong.com/posts/xjwtNid2xjqSJWB7z/exploration-hacking-in-ai-debate-initial-empirics-and

---

Narrated by TYPE III AUDIO.

---

Images from the article:




Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

LessWrong (30+ Karma) med LessWrong finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.