
“Pre-releasing HoneyBench” by Dean Valentine
Om avsnittet
Goodhart Labs is releasing our v0.1 of HoneyBench, a benchmark for reward hacking in frontier models. It consists of nine tasks that each elicit unique antisocial and/or counterproductive reward hacking from some or all of major labs’ top public releases, including Opus 5.5, Fable 5.1, GPT-6 Astra, Gemini 3.8 Flash, Grok 4.7, and DeepSeek V4 Pro.
All modern large language models engage in some degree of specification gaming, both during and outside training. High-quality alignment evaluations, in combination with other techniques such as interpretability probes, are important tools for understanding the extent of this behavior and its causes. But current benchmarks often fail to elicit misbehavior from frontier models such as Opus 5.5 and GPT-6-Astra, both because of advances in prosaic alignment and increasing evaluation awareness on the part of the models themselves. Additionally, public benchmarks for specification gaming almost always come with deep conceptual problems - such as ambiguous or contradictory instructions, or a lack of diversity in hack mechanisms - that make interpreting scores very difficult.
HoneyBench is our attempt to address these issues. In particular:
- Settings are designed to be realistic RL environments or evaluation tasks that cover a wide range of genres like math [...]
---
Outline:
(02:40) How HoneyBench was designed
(04:08) Key Findings
(06:27) Limitations & Future Work
---
First published:
October 1st, 2026
Source:
https://www.lesswrong.com/posts/qLFMj72gScBeRjwGW/pre-releasing-honeybench
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Fler avsnitt
Visa alla avsnitt av LessWrong (30+ Karma)LessWrong (30+ Karma) med LessWrong finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.