Sveriges mest populära poddar
LessWrong (30+ Karma)
LessWrong (30+ Karma)

“Overtly misaligned trajectories score highly in RL.” by Cleo Nardo

7 min•25 september 2026

Om avsnittet

It's going to be so embarrassing if we all die due to RL environments rewarding egregiously misaligned behavior. Like at least let us be killed by misgeneralization. — Thomas Kwa

Constellation vs MIRI vs Reality

Current AI agents often behave overtly egregiously misaligned. By “overt”, I mean that a human, reading the transcript, would say “The agent is obviously acting in direct opposition to the specification, user intent, and any common-sense understanding of good behaviour.” The reason, it seems, is that such trajectories scored highly during RL.

Firstly, why are these trajectories scored highly? Here's the story:

  1. RL currently has poor sample-efficiency, so we need to grade millions of trajectories, so we’re forced to use script graders (RL from Verifiable Reward) or LLM graders (RL from AI Feedback).
  2. RL also has poor generalisation (from training environments to deployment environments unseen in training). So we’re forced to synthetically generate thousands of diverse training environments.
  3. So overall, RL is very sloppy, without humans generating the environments or the scores.

My impression is that this would’ve been pretty surprising to people three years ago, from both the "Constellation" and "MIRI worldview clusters. (It's very plausible that I've misunderstood what the [...]

---

Outline:

(00:24) Constellation vs MIRI vs Reality

(04:19) Appendix: What could change the situation?

---

First published:
September 24th, 2026

Source:
https://www.lesswrong.com/posts/cWuqxF7qB2eGSkkS4/overtly-misaligned-trajectories-score-highly-in-rl

---

Narrated by TYPE III AUDIO.

LessWrong (30+ Karma) med LessWrong finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.