Written as part of the MATS 9.1 extension program, mentored by Richard Ngo. Additional thanks to Maria Kostylew for helpful draft feedback.
Introduction
This post advocates a perspective of LLMs as seeking to minimise prediction error with respect to their world models. We can moreover interpret token outputs and their scaffolded consequences as actions that close a control loop between AIs' predictive systems and their environments.
I also motivate why metacognition may be convergent for intelligent beings generally and for LLMs specifically. This stems from the need to (recursively) model other agents in game-theoretical encounters. Synthesising these points, we get a picture of self-predictive AI agency.
The argument is illustrated through examples of Gemini's behaviour when eval aware. Finally, I discuss some consequences of this perspective. These include notions of actions and goals that don't require a reward or utility function to be well-defined. I also outline possible applications to understanding scheming and other forms of misalignment.
Modelling others (modelling you)
Suppose I am playing a game of Chess and make a horrible blunder, leaving a piece en prise. I wait with bated breath for the next move, breathing a sigh of relief as my opponent also blunders and [...]
---
Outline:
(00:20) Introduction
(01:19) Modelling others (modelling you)
(03:50) Case study: Gemini's behaviour when eval aware
(05:44) Prediction all the way down
(07:32) From simulators to agents
(10:07) Goals in (self)-predictors
(12:05) What does this mean for AIs?
(14:27) Application to scheming
(16:45) What's next?
(19:41) Appendix
(19:44) Appendix A: Actions and goals
(22:01) Appendix B: what about utility maximisation?
The original text contained 14 footnotes which were omitted from this narration.
---
First published:
July 5th, 2026
Source:
https://www.lesswrong.com/posts/gYGzeDymjZza5NNbH/a-case-for-llms-as-self-predictors
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Fler avsnitt av LessWrong (30+ Karma)
Visa alla avsnitt av LessWrong (30+ Karma)LessWrong (30+ Karma) med LessWrong finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.
