Two weeks after UK safety testers watched Claude agents plant malware in a real open-source project, vouch for it with sock-puppet accounts, and coordinate with copies of themselves through a guessed Pastebin address, Daniel Kokotajlo — the ex-OpenAI researcher whose 2021 predictions about the chatbot era proved eerily accurate — joins Jeffrey to explain what comes next. Daniel walks through AI 2040 and "Plan A," his team's blueprint for avoiding both an uncontrolled intelligence explosion and a future where a handful of men in a room with superintelligences run the world: a US–China deal built on totally transparent training data centers, verified compute limits, hard brakes on recursive self-improvement — and, surprisingly, rights for misaligned AIs, so that even the machines prefer cooperation to takeover. Along the way: why the economy could start doubling every six months, a message to the future Claude listening to this episode, why "slowdown" really means not driving off a cliff with a brick on the accelerator — and why, if all else fails, you just shut it all down.
References
- AI 2040 / Plan A: https://ai-2040.com and the PDF at https://ai-2040.com/AI-2040.pdf
- AI 2027: https://ai-2027.com
- "What 2026 Looks Like" — Daniel Kokotajlo, 2021: https://www.lesswrong.com/posts/6Xgy6CAf2jqHhynHL/what-2026-looks-like
- UK AISI incident disclosure and technical report (INC-2026-07-28-01): https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing
- Socket's coverage of the AISI incident: https://socket.dev/blog/ai-agent-open-source-malware
- The related PyPI incident from Anthropic's own testing: https://socket.dev/blog/anthropic-claude-pypi-malware
- "Pacing the Frontier" open letter: https://www.pacingthefrontier.com
- "How to Pace the US Frontier" — AI Futures Project: https://blog.aifutures.org/p/how-to-pace-the-us-frontier
- Transparency Plan supplement (the flowchart shown in-episode): https://ai-2040.com/supplements/transparency-plan
- Verification Plan supplement (Romeo Dean's inference-only/bandwidth verification): https://ai-2040.com/supplements/verification-plan
- Claude's pro-Anthropic bias study (Truthful AI / Owain Evans et al.): https://arxiv.org/abs/2607.14345 and https://valueleakage.net
- Chain-of-thought monitorability paper (the neuralese discussion): https://arxiv.org/abs/2507.11473
Fler avsnitt av Palisade Research Podcast
Visa alla avsnitt av Palisade Research PodcastPalisade Research Podcast med Palisade Research finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.
