This is the fifth post in the sequence Implications of Continual Learning for LLM Agents.
Summary
While writing our continual learning sequence, we sent a survey to a number of AI safety researchers with questions about continual learning. This post summarizes the results of that survey. We asked whether respondents agree with various arguments we advance throughout the sequence, how worried respondents are about certain risks, how respondents would forecast different aspects of the future of CL, and how promising respondents find various proposed angles of attack. We also asked open-ended questions about the benefits of CL and whether we seem to be missing any major considerations. At the end of the post, we also provide an overview of forecasts about CL made by other experts who didn’t participate in our survey.
We received survey responses from:
- Ryan Faulkner, PhD student at the University of Toronto focusing on multi-agent simulation, learning, and cooperation
- Nikola Jurkovic, Member of Technical Staff at METR
- Alex Mallen, Member of Technical Staff at Redwood Research, doing research and writing on AI threat models. Author of "The case for countermeasures to memetic spread of misaligned values"
- Evgenii Opryshko, 3rd year PhD student at the [...]
---
Outline:
(00:20) Summary
(02:20) Broad takeaways
(03:59) Full results
(04:56) Futures
(05:58) Reflection and goal drift
(07:25) Loss of the last-mover advantage
(08:24) Control
(09:48) Angles of attack
(11:51) Open-ended questions
(13:44) Forecasts from other experts
(14:02) AI 2027
(14:57) IABIED
(15:43) Understanding AI Trajectories: Mapping the Limitations of Current AI Systems
(16:17) Brain-like AGI safety
(18:20) Other forecasts and opinions
---
First published:
June 24th, 2026
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Fler avsnitt av LessWrong (30+ Karma)
Visa alla avsnitt av LessWrong (30+ Karma)LessWrong (30+ Karma) med LessWrong finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.
