Sveriges mest populära poddar
Redwood Research Blog

“Training-time schemers vs behavioral schemers” by Alex Mallen

13 min6 maj 2025

Subtitle: Clarifying ways in which faking alignment during training is neither necessary nor sufficient for the kind of scheming that AI control tries to defend against..

People use the word “schemer” in two main ways:

  1. “Scheming” (or similar concepts: “deceptive alignment”, “alignment faking”) is often defined as a property of reasoning at training-time1. For example, Carlsmith defines a schemer as a power-motivated instrumental training-gamer—an AI that, while being trained, games the training process to gain future power. I’ll call these training-time schemers.

  2. On the other hand, we ultimately care about the AI's behavior throughout the entire deployment, not its training-time reasoning, because, in order to present risk, the AI must at some point not act aligned. I’ll refer to AIs that perform well in training but eventually take long-term power-seeking misaligned actions as behavioral schemers. When people say that a model is a schemer and [...]

---

Outline:

(02:23) Training-time schemers that continue to act aligned for the whole deployment

(02:42) Empirical evidence

(05:19) Theoretical argument

(08:01) Related ambiguity about situational awareness

(09:12) Behavioral schemers that weren't training-time schemers

(12:00) Discussion

---

First published:
May 6th, 2025

Source:
https://redwoodresearch.substack.com/p/training-time-schemers-vs-behavioral

---

Narrated by TYPE III AUDIO.

---

Images from the article:

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Fler avsnitt av Redwood Research Blog

Visa alla avsnitt av Redwood Research Blog

Redwood Research Blog med Redwood Research finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.