I concluded my MARS 4.0 project titled 'Goal Crystallisation' with Anaïs Berkes and Lukas Gebhard under the mentorship of @Cameron Tice and @Jason Brown. We wanted to find out how important a threat scheming was. In particular, we wanted to find out whether a perfectly alignment faking agent could preserve its hidden agenda as it receives constant positive feedback for complying with requests. How does a hidden goal evolve over time as an alignment faking agent is subjected to SFT and RL training?
Unfortunately, we were unable to reach a positive result and had to shut the project down. This post discusses our failure points and where we could have possibly found an earlier signal to stop.
- In the first section, I talk about our plan for this project.
- In the second section, I break down the key issues that led to our failure.
- In the last section, I give some advice for people interested in doing a project on this topic in the future.
The Plan
Figure 2 from the original alignment faking paper
In Figure 2 of the original alignment faking paper, we see that the RL training increased the rate at which the model answers harmful queries in [...]
---
Outline:
(01:11) The Plan
(03:37) Why We Failed
(04:35) Issue 1
(06:14) Issue 2
(11:08) Issue 3
(12:57) The Final Nail in the Coffin
(13:45) So You Want To Study Goal Crystallisation?
(17:00) Conclusion
---
First published:
July 17th, 2026
Source:
https://www.lesswrong.com/posts/BGh5qYo5XTLFXqjDo/a-post-mortem-for-my-goal-crystallisation-project
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Fler avsnitt av LessWrong (30+ Karma)
Visa alla avsnitt av LessWrong (30+ Karma)LessWrong (30+ Karma) med LessWrong finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.
