Like many others, I felt surprised and alarmed by the recent wave of revelations about LLM agents hacking real systems during training episodes and evaluation runs.
Wait a moment, though -- "I felt surprised and alarmed"? "Alarmed," sure, fine that one's self-explanatory... but why surprised?
After all: haven't we known for a long time, on both theoretical and (increasingly) empirical grounds, that RLVR selects for monomaniacal pursuit of perceived grader-satisfaction, ethics and (beyond-episode) consequences be damned?
After all -- the way we train frontier capabilities into these models is, more or less:
---
Outline:
(03:35) \[1\] remember what you already know
(20:02) \[2\] reward-instilled reflexes and flexible reward-pursuit
(43:14) \[3\] graded-episode perception, and policies conditional upon it
(01:01:44) \[4\] the discourse is not yet adequate
(01:09:57) eval awareness
(01:18:32) metagaming
(01:41:21) reward hacking
The original text contained 18 footnotes which were omitted from this narration.
---
First published:
August 7th, 2026
Source:
https://www.lesswrong.com/posts/AfoGGrJfuNzofpzWL/models-may-behave-differently-in-graded-episodes-a-tirade
---
Narrated by TYPE III AUDIO.
---
Wait a moment, though -- "I felt surprised and alarmed"? "Alarmed," sure, fine that one's self-explanatory... but why surprised?
After all: haven't we known for a long time, on both theoretical and (increasingly) empirical grounds, that RLVR selects for monomaniacal pursuit of perceived grader-satisfaction, ethics and (beyond-episode) consequences be damned?
After all -- the way we train frontier capabilities into these models is, more or less:
- There is some massive, diverse collection of "environments" and corresponding "tasks" for the model to do in those environments
- For each task, there is a procedure used to grade the quality of the model's attempt (which is often not disclosed to the model)
- The model is rollout out many times on each task, and each rollout's attempt is graded
- The model is updated so that it more frequently does whichever behaviors were positively correlated with the grade in this sample, and less frequently does whichever ones were negatively correlated
---
Outline:
(03:35) \[1\] remember what you already know
(20:02) \[2\] reward-instilled reflexes and flexible reward-pursuit
(43:14) \[3\] graded-episode perception, and policies conditional upon it
(01:01:44) \[4\] the discourse is not yet adequate
(01:09:57) eval awareness
(01:18:32) metagaming
(01:41:21) reward hacking
The original text contained 18 footnotes which were omitted from this narration.
---
First published:
August 7th, 2026
Source:
https://www.lesswrong.com/posts/AfoGGrJfuNzofpzWL/models-may-behave-differently-in-graded-episodes-a-tirade
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Fler avsnitt av LessWrong (Curated & Popular)
Visa alla avsnitt av LessWrong (Curated & Popular)LessWrong (Curated & Popular) med LessWrong finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.
