Sveriges mest populära poddar
LessWrong (30+ Karma)

“A case for LLMs as Self-predictors” by Ashe Vazquez Nuñez

23 min5 juli 2026

Written as part of the MATS 9.1 extension program, mentored by Richard Ngo. Additional thanks to Maria Kostylew for helpful draft feedback.

Introduction

This post advocates a perspective of LLMs as seeking to minimise prediction error with respect to their world models. We can moreover interpret token outputs and their scaffolded consequences as actions that close a control loop between AIs' predictive systems and their environments.

I also motivate why metacognition may be convergent for intelligent beings generally and for LLMs specifically. This stems from the need to (recursively) model other agents in game-theoretical encounters. Synthesising these points, we get a picture of self-predictive AI agency.

The argument is illustrated through examples of Gemini's behaviour when eval aware. Finally, I discuss some consequences of this perspective. These include notions of actions and goals that don't require a reward or utility function to be well-defined. I also outline possible applications to understanding scheming and other forms of misalignment.

Modelling others (modelling you)

Suppose I am playing a game of Chess and make a horrible blunder, leaving a piece en prise. I wait with bated breath for the next move, breathing a sigh of relief as my opponent also blunders and [...]

---

Outline:

(00:20) Introduction

(01:19) Modelling others (modelling you)

(03:50) Case study: Gemini's behaviour when eval aware

(05:44) Prediction all the way down

(07:32) From simulators to agents

(10:07) Goals in (self)-predictors

(12:05) What does this mean for AIs?

(14:27) Application to scheming

(16:45) What's next?

(19:41) Appendix

(19:44) Appendix A: Actions and goals

(22:01) Appendix B: what about utility maximisation?

The original text contained 14 footnotes which were omitted from this narration.

---

First published:
July 5th, 2026

Source:
https://www.lesswrong.com/posts/gYGzeDymjZza5NNbH/a-case-for-llms-as-self-predictors

---

Narrated by TYPE III AUDIO.

---

Images from the article:


Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Fler avsnitt av LessWrong (30+ Karma)

Visa alla avsnitt av LessWrong (30+ Karma)

LessWrong (30+ Karma) med LessWrong finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.