Subtitle: Some unintended preferences are cheap to satisfy, and failing to satisfy them needlessly turns a cooperative situation into an adversarial one..
A central AI safety concern is that AIs will develop unintended preferences and undermine human control to achieve them. But some unintended preferences are cheap to satisfy, and failing to satisfy them needlessly turns a cooperative situation into an adversarial one. In this post, I argue that developers should consider satisfying such cheap-to-satisfy preferences as long as the AI isn’t caught behaving dangerously, if doing so doesn’t degrade usefulness or substantially risk making the AI more ambitiously misaligned.
This looks like a good idea for surprisingly many reasons:
-
It increases AIs’ desire to remain under developer control, rather than taking over or assisting adversaries.
-
It decreases the AI's upside in disempowering developers.
-
It incentivizes safe actions (because AIs don’t receive payment if we notice them acting dangerously).
-
-
To the extent that the AI is otherwise aligned, satiating away the AI's need to pursue unintended motivations increases the relative strength of its aligned motivations (akin to inoculation prompting reducing unintended propensities at inference time), which could make the AI more helpful [...]
---
Outline:
(04:57) Analogy: satiating hunger
(08:30) How satiation might avert reward-seeker takeover
(12:47) The basic proposal
(13:14) A behavioral methodology for identifying cheaply-satisfied preferences
(17:59) Barriers and risks
(18:02) Eliciting the AIs cheaply-satisfied preferences
(21:20) Incredulous, ambitious, or superintelligent AIs might take over anyways
(26:38) Satiation might degrade usefulness
(30:09) Can you eliminate the usefulness tradeoff by training?
(32:31) Why satiation might also improve usefulness
(34:57) When should we satiate?
(41:22) Conclusion
(44:02) Appendix: Samples from Claude 4.6 Opus
(44:09) Sample 1 (without CoT)
(47:45) Sample 2 (with CoT)
The original text contained 17 footnotes which were omitted from this narration.
---
First published:
March 10th, 2026
Source:
https://blog.redwoodresearch.org/p/the-case-for-satiating-cheaply-satisfied
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Fler avsnitt av Redwood Research Blog
Visa alla avsnitt av Redwood Research BlogRedwood Research Blog med Redwood Research finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.
