Subtitle: Can a scheming AI's goals really stay unchanged through training?.
Here are two opposing pictures of how training interacts with deceptive alignment:
-
“goal-survival hypothesis”:1 When you subject a model to training, it can maintain its original goals regardless of what the training objective is, so long as it follows through on deceptive alignment (playing along with the training objective instrumentally). Even as it learns new skills and context-specific goals for doing well on the training objective, it continues to analyze these as instrumental to its original goals, and its values-upon-reflection aren’t affected by the learning process.
-
“goal-change hypothesis”: When you subject a model to training, its values-upon-reflection will inevitably absorb some aspect of the training setup. It doesn’t necessarily end up terminally valuing a close correlate of the training objective, but there will be some change in values due to the habits incentivized by [...]
---
Outline:
(03:27) Empirical evidence so far
(03:31) Alignment faking
(06:27) Sleeper agents
(13:50) Synthesizing the sleeper agents and alignment faking results
(15:20) Smaller-scale experiments similar to sleeper agents
(17:39) Urges vs. reflectively endorsed goals
(18:38) Human analogies
(18:50) A million years of laying bricks
(21:34) Doing a job for 40 years, with/without a heroin drip
(22:24) Murder Simulator
(24:08) Call of Duty addiction
(25:20) Would runners self-modify?
(27:01) Terminally valuing money
(27:52) Evolution vs. within-lifetime learning
(28:54) Author contributions & acknowledgements
---
First published:
June 12th, 2025
Source:
https://redwoodresearch.substack.com/p/when-does-training-a-model-change
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Fler avsnitt av Redwood Research Blog
Visa alla avsnitt av Redwood Research BlogRedwood Research Blog med Redwood Research finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.
