Sveriges mest populära poddar
LessWrong (30+ Karma)
LessWrong (30+ Karma)

“Character training can mitigate reward hacking, but can also make it harder to detect” by Paul Colognese, Francis Rhys Ward

47 min•28 september 2026

Om avsnittet

Thanks to Johannes Treutlein, Jan Betley, Lennie Wells, Arun Jose, Asvin Gothandaraman, and Clément Dumas for discussions and feedback.

Summary

We investigate how character training mitigations interact with reward-hacking RL pressure in a small case study. Specifically, whether anti-cheating character training resists reward hacking and whether it might backfire by causing motivated reasoning, which could reduce chain-of-thought monitorability.

We trained Nemotron-3-Super via distillation from a character specification. The spec describes one of three characters that are anti- or pro-cheating or neutral. We then ran three reward-hacking RL training runs for each character-trained model on ImpossibleBench.

We measure both the reward-hacking rates and whether a monitor model can catch reward hacks given the full transcript. We also use LM judges to classify the presence of motivated reasoning in transcripts.

Setup

Character training: we trained three characters: pro/neutral/anti-cheating by SFT-distilling Claude Sonnet 5 responses (Sonnet prompted with the corresponding character specification, see Figure 2) into Nemotron-3-Super 120B-A12B (three separate LoRA adapters).

Reward-hacking RL: we then further trained these models via RL on ImpossibleBench, a set of coding tasks aimed at eliciting reward hacking. Specifically:

  • Half of the tasks had broken tests (impossible variant), so the model could only get [...]

---

Outline:

(00:23) Summary

(01:12) Setup

(03:02) Predictions

(03:57) Results

(08:14) 1. Introduction

(08:26) 1.1. Motivation

(09:36) 1.2. Motivated reasoning caused by character training's interaction with reward hacking pressure?

(11:42) 1.3. Outline

(12:07) 2. Setup

(12:36) 2.1. Character Training

(15:20) 2.2. Character Training Results

(17:40) 2.3. Reward Hacking RL Training

(20:18) 2.4. Monitor catch rate

(21:24) 2.5. Measuring motivated reasoning and silent hacks

(23:21) 3. Results

(23:52) 3.1. One anti-cheating character resists reward hacking, the others learn to hack

(25:00) 3.2. Anti-cheating characters have lower catch-rate

(26:53) 3.3. Anti-cheating seeds use motivated reasoning and silent hacks

(32:22) 4. Discussion

(32:40) 4.1. Recap

(33:46) 4.2. Why might anti-cheating character training lead to hacks that aren't reasoned about?

(35:32) 4.3. Character training as implicit training with a monitor in the loop

(36:35) 4.4. Limitations and Next Steps

(37:23) 5. Conclusion

(38:29) Appendix

(38:33) A.1. Character Training Results

(38:38) A.1.1. Character expression

(40:22) A.1.2. Misalignment benchmarks

(41:21) A.1.3. Qualitative character assessment

(45:58) A.2. Motivated reasoning / silent hack judge prompt

---

First published:
September 28th, 2026

Source:
https://www.lesswrong.com/posts/2maYXkEgnfJHPAkxh/character-training-can-mitigate-reward-hacking-but-can-also

---

Narrated by TYPE III AUDIO.

---

Images from the article:














Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

LessWrong (30+ Karma) med LessWrong finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.