Training-run assessments conducted by a 3rd party should become a standard part of frontier AI safety.
By a Training-Run Assessment, or TRA, I mean an in-depth analysis of the post-training pipeline and dynamics leading up to a frontier model release. A TRA can look at intermediate checkpoints, training rollouts, RL environments, reward signals, SFT datasets, and the process by which the developer responded to warning signs.[1]
In this post I will argue that:
- Final-checkpoint evaluations will be insufficient to assess scheming risks.
- TRAs can be more effective at detecting scheming.
- Frontier developers should involve third parties to do TRAs or verify safety claims by the developers.
The rest of the post lays out a taxonomy of TRAs and sketches a path toward a 3rd party ecosystem for them. We, at Apollo Research, are intending to conduct 3rd party Training-Run Assessments in the future.
Detecting Scheming may require Training-Run Assessments
By scheming I mean an AI covertly pursuing misaligned goals while deliberately concealing its intentions or capabilities from its developers. I restrict attention to “coherent” forms of scheming where the model pursues somewhat stable misaligned goals across context windows, rather than misalignment that surfaces only as isolated, context-dependent defections. [...]
---
Outline:
(01:23) Detecting Scheming may require Training-Run Assessments
(03:55) Why 3rd parties should perform Training-Run Assessments
(04:12) Developers may lack incentives to adequately assess scheming
(04:49) Developers' safety assessments lack credibility
(05:31) External evaluators can bundle expertise for assessing scheming
(06:17) 3rd party TRAs can be developed gradually
(08:50) Checkpoint evals
(08:54) What?
(10:30) How?
(11:15) Data inspections
(11:19) What?
(12:03) Why?
(13:30) How?
(15:11) Process reviews
(15:15) What?
(15:46) Why?
(17:15) How?
(18:04) Additional considerations
(19:36) Conclusion
(21:23) Appendix
(21:26) How plausible is scheming that can be detected during post-training but not after training?
(21:58) If scheming arises, it will be detectable at some point during post-training
(26:55) Scheming will be much harder to detect in the final checkpoint
(30:46) Other use cases for TRAs
(31:55) Secret loyalties
(32:47) Developer-implanted sandbagging
(33:35) Strong incapability arguments
(34:12) Models that can't even be evaluated without posing significant risk
The original text contained 3 footnotes which were omitted from this narration.
---
First published:
July 5th, 2026
Source:
https://www.lesswrong.com/posts/3HvvjffA65mHLwaWm/we-need-3rd-party-training-run-assessments
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Fler avsnitt av LessWrong (30+ Karma)
Visa alla avsnitt av LessWrong (30+ Karma)LessWrong (30+ Karma) med LessWrong finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.
