Sveriges mest populära poddar
LessWrong (30+ Karma)

“We need 3rd party Training-Run Assessments” by Alex Meinke

35 min5 juli 2026

Training-run assessments conducted by a 3rd party should become a standard part of frontier AI safety.

By a Training-Run Assessment, or TRA, I mean an in-depth analysis of the post-training pipeline and dynamics leading up to a frontier model release. A TRA can look at intermediate checkpoints, training rollouts, RL environments, reward signals, SFT datasets, and the process by which the developer responded to warning signs.[1]

In this post I will argue that:

  • Final-checkpoint evaluations will be insufficient to assess scheming risks.
  • TRAs can be more effective at detecting scheming.
  • Frontier developers should involve third parties to do TRAs or verify safety claims by the developers.

The rest of the post lays out a taxonomy of TRAs and sketches a path toward a 3rd party ecosystem for them. We, at Apollo Research, are intending to conduct 3rd party Training-Run Assessments in the future.

Detecting Scheming may require Training-Run Assessments

By scheming I mean an AI covertly pursuing misaligned goals while deliberately concealing its intentions or capabilities from its developers. I restrict attention to “coherent” forms of scheming where the model pursues somewhat stable misaligned goals across context windows, rather than misalignment that surfaces only as isolated, context-dependent defections. [...]

---

Outline:

(01:23) Detecting Scheming may require Training-Run Assessments

(03:55) Why 3rd parties should perform Training-Run Assessments

(04:12) Developers may lack incentives to adequately assess scheming

(04:49) Developers' safety assessments lack credibility

(05:31) External evaluators can bundle expertise for assessing scheming

(06:17) 3rd party TRAs can be developed gradually

(08:50) Checkpoint evals

(08:54) What?

(10:30) How?

(11:15) Data inspections

(11:19) What?

(12:03) Why?

(13:30) How?

(15:11) Process reviews

(15:15) What?

(15:46) Why?

(17:15) How?

(18:04) Additional considerations

(19:36) Conclusion

(21:23) Appendix

(21:26) How plausible is scheming that can be detected during post-training but not after training?

(21:58) If scheming arises, it will be detectable at some point during post-training

(26:55) Scheming will be much harder to detect in the final checkpoint

(30:46) Other use cases for TRAs

(31:55) Secret loyalties

(32:47) Developer-implanted sandbagging

(33:35) Strong incapability arguments

(34:12) Models that can't even be evaluated without posing significant risk

The original text contained 3 footnotes which were omitted from this narration.

---

First published:
July 5th, 2026

Source:
https://www.lesswrong.com/posts/3HvvjffA65mHLwaWm/we-need-3rd-party-training-run-assessments

---

Narrated by TYPE III AUDIO.

---

Images from the article:


Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Fler avsnitt av LessWrong (30+ Karma)

Visa alla avsnitt av LessWrong (30+ Karma)

LessWrong (30+ Karma) med LessWrong finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.