
“Continual learning might make your blocking monitors nearly useless” by Alex Mallen
Om avsnittet
Subtitle: When monitor evasion looks like legitimate learning to your continual learning system, it's hard to have one without the other.
Many control protocols work by intervening on an untrusted AI's actions during deployment. For example, you might set up a monitor that scores each action's suspiciousness and blocks actions above a threshold, replacing them with actions from a weaker “trusted” model (a defer-to-trusted protocol). Interventions like this cost usefulness: the protocol sometimes replaces good actions with worse ones.
Continual learning, by which I mean any mechanism that lets an AI accumulate skills and knowledge over the course of its deployment, optimizes for usefulness. So, if blocking monitors get in the way of task success, I expect continually-learning AIs to learn how to evade blocking monitors. This requires no scheming; it falls out of usefulness pressure applied to a benign model. And the effect is plausibly substantial: a long deployment with online RL (for example) amounts to training the policy against the monitor, and could involve enough optimization to leave blocking monitors nearly useless.
While the problem would be noticeable to developers, I argue it is hard to fix. Because monitor evasion looks like legitimate learning to [...]
---
Outline:
(02:35) Online RL trains the policy against your monitors
(05:20) The argument is weaker for memory-based continual learning
(07:21) The argument doesn't require that adversarial models will exploit continual learning
(08:28) Easy to notice, hard to fix
(09:27) Conservative control evaluations don't fix blocking monitors
(10:54) What developers could do
(10:57) You can't just not use blocking monitors
(11:45) But you can try to reduce the usefulness cost of blocking protocols
(12:48) You can't just hold out blocked trajectories from training
(16:47) But you can try to improve the monitor online
(19:39) Conclusion
(20:39) Appendix: How memory-based systems might avoid this issue, and how they might fail
The original text contained 16 footnotes which were omitted from this narration.
---
First published:
September 25th, 2026
Source:
https://blog.redwoodresearch.org/p/continual-learning-might-make-your
---
Narrated by TYPE III AUDIO.
Fler avsnitt
Visa alla avsnitt av Redwood Research BlogRedwood Research Blog med Redwood Research finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.