Sveriges mest populära poddar
Redwood Research Blog

“Reward-seekers will probably behave according to causal decision theory” by Alex Mallen

3 min28 mars 2026

Subtitle: They'd renege on non-binding commitments, defect against copies of themselves in prisoner's dilemmas, etc..

Background: There are existing arguments to the effect that default RL algorithms encourage CDT reward-maximizing behavior on the training distribution. (That is: Most RL algorithms search for policies by selecting for actions that cause high reward. E.g., in the twin prisoner's dilemma, RL algorithms randomize actions conditional on the policy, which means that the action provides no evidence to the RL algorithm about the counterparty's action.1) This doesn’t imply RL produces CDT reward-maximizing policies: CDT behavior on the training distribution doesn’t imply CDT generalization because agents can fake CDT in the same way that they can fake alignment, or might develop arbitrary other propensities that were correlated with reward on the training distribution.

But conditional on reward-on-the-episode seeking, the AI is likely to generalize CDT.

If, for example, a reward-seeker tried to evidentially cooperate between episodes (so it had non-zero regard for reward that isn’t used to reinforce its current actions), this would be trained away because the AI would be willing to give up reward on the current episode to some extent. You might be tempted to respond with: “But can’t the [...]

The original text contained 2 footnotes which were omitted from this narration.

---

First published:
March 28th, 2026

Source:
https://blog.redwoodresearch.org/p/reward-seekers-will-probably-behave

---

Narrated by TYPE III AUDIO.

Fler avsnitt av Redwood Research Blog

Visa alla avsnitt av Redwood Research Blog

Redwood Research Blog med Redwood Research finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.