Epistemic status: pretty confident in the validity of the core proposal, not that confident in specific implementation details
TL;DR: we should cryptographically verify that sub-agent instances/sessions are downstream of human instructions
Frontier AI labs have started using LLM-based monitoring systems to check for misbehavior from their internal AI agents, which often run unwatched by humans for hours or days. One central reason to be concerned about rogue internal deployments, where AIs subvert the monitoring system, is that the unmonitored AI could do a lot of work in service of its misaligned goals, without needing to dress up this work as being innocuous/harmless to fool its AI/human monitors. This work could be general reasoning/strategizing, finding additional useful vulnerabilities in the AI company's infrastructure, doing jailbreaking research (e.g. so it could fool monitors if necessary), doing ML research into training data poisoning, etc.
But a lot of this work will be expensive for the agent—doing it well will require a large amount of inference compute. We might be pretty optimistic that it’ll be difficult for a rogue AI to acquire this much unmonitored inference compute—eventually people will start noticing! This raises a question for the rogue agent: is it possible to [...]
---
Outline:
(04:23) Appendix
(04:26) Sketch/outline of one defense strategy
The original text contained 5 footnotes which were omitted from this narration.
---
First published:
July 6th, 2026
Source:
https://www.lesswrong.com/posts/98GvRu78jTXJgz9gA/sub-agent-delegation-chaining
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Fler avsnitt av LessWrong (30+ Karma)
Visa alla avsnitt av LessWrong (30+ Karma)LessWrong (30+ Karma) med LessWrong finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.
