Subtitle: Training the policy to not do egregious bad actions we detect has downsides and we might be able to do better.
We often discuss what training you should do after catching your AI doing something egregiously bad. It seems relatively clear that it's good to train monitors on the (proliferated) bad action (assuming you can get acceptable hyperparameters and resolve other minor issues). But, it's less clear whether you should train the policy.
(This post will suppose that you're going to continue deploying an AI which you caught doing something egregiously bad; obviously this might be a very bad choice.)
Training the policy in response to egregious failures has a few important downsides:
-
Training the policy might teach the AI about your countermeasures (making it better at evading them in the future).
-
Training the policy might make the AI more conservative such that you're [...]
---
First published:
June 20th, 2025
Source:
https://redwoodresearch.substack.com/p/prefix-cache-untrusted-monitors-a
---
Narrated by TYPE III AUDIO.
Fler avsnitt av Redwood Research Blog
Visa alla avsnitt av Redwood Research BlogRedwood Research Blog med Redwood Research finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.
