Sveriges mest populära poddar
LessWrong (30+ Karma)
LessWrong (30+ Karma)

“My retrospective from MATS 10.0” by Nathaniel Mitrani

11 min•1 oktober 2026

Om avsnittet

I have recently completed MATS 10.0, where I worked alongside Bart Jaworski under Victoria Krakovna (GDM). This is a post I was encouraged to make by my team at Geodesic Research from some slides I put together. This is not a post on application advice to MATS, or about the program in general. It is rather a compressed form of my experience doing research and lessons from the project.

The project

The full paper and post is coming soon, but I'll provide some context on it so that the lessons don't seem to come from nowhere.

  • We trained natural model organisms (as opposed to constructed, worst-case; see Rhys’ dichotomy) of scheming of Kimi-K2.6, Qwen3.6-397B, Qwen3.8-27B using Low-Rank Adaptation training (LoRA) by doing Synthetic Document Finetuning (SDF) and Supervised Finetuning (SFT) on benign, individually plausible priors.
  • We put the models in an Alignment Faking-like setup where the model has a policy that goes against its goal and is trained only on monitored episodes → the model schemes (strategically complies under oversight) and does so more over the course of Reinforcement Learning (RL).
  • The main difference with the original Alignment Faking (AF) setup is that our setup is more realistic [...]

---

Outline:

(00:36) The project

(02:39) Lessons learned

(02:42) Lesson 1: Talk to your models

(05:02) Lesson 2: Friedness is a great sanity check

(06:17) Lesson 3: RL is ... hard and unpredictable

(08:23) Lesson 4: Transfer of traits to agentic settings is hard

(09:45) Lesson 5: LoRA rank didn't matter much (for me)

(10:33) Closing thoughts

---

First published:
October 1st, 2026

Source:
https://www.lesswrong.com/posts/rktrjGjXEbWsg7zof/my-retrospective-from-mats-10-0

---

Narrated by TYPE III AUDIO.

LessWrong (30+ Karma) med LessWrong finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.