In our last post, we argued that measuring evaluation awareness is fundamentally challenging because of the safe-to-dangerous distributional shift: we cannot directly measure the evaluation awareness of a model without deploying it, but we cannot safely deploy it until we know it is not scheming. We expect sufficiently superhuman AI will be eval aware, but this post outlines a tentative solution for bounding the awareness of a ~human-expert-level[1] AI across this safe-to-dangerous shift:
Instead of letting the untrusted model probe its environment to discriminate eval from deployment—and risk harmful actions—we can use a trusted human red team augmented with LLMs. If the human red team cannot differentiate between the environments, then we have some evidence by analogy that expert-level LLMs also could not differentiate between them.
In this post we:
- (S1) Motivate robust upper bounds of evaluation awareness as an important but insufficient step to ruling out scheming in human-expert-level AI.
- (S2) Explain the core of our idea for bounds using a human red-teaming experiment.
- (S3) Propose to address superhuman LLM Truesight by augmenting humans with safely SFTed LLMs.
- (S4) Propose to reduce costs of human labour by starting with cheaper experiments to rule out obvious evaluations.
- (S5) Explore whether [...]
---
Outline:
(01:44) 1. Why upper-bound eval awareness of expert-level AI?
(03:44) 2. Proposal: Evaluate evaluations with human red teams
(05:47) 3. Problem 1: Can the red team upper-bound the LLM's capabilities?
(08:55) 4. Problem 2: How can we reduce the cost?
(10:43) 5. Problem 3: Is the red team's challenge analogous enough to the LLM's?
(11:05) How close is close enough?
(12:31) How close is the proxy?
(14:36) 6. Conclusion
The original text contained 3 footnotes which were omitted from this narration.
---
First published:
July 6th, 2026
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Fler avsnitt av LessWrong (30+ Karma)
Visa alla avsnitt av LessWrong (30+ Karma)LessWrong (30+ Karma) med LessWrong finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.
