Subtitle: The complement to control evaluations.
One way to show that risk from deploying an AI system is small is by showing that the model is not capable of subverting security measures enough to cause substantial harm. In AI control research so far, control evaluations have measured whether a red-team-created attack policy can defeat control measures, but they haven't measured how effectively models can generate such attack policies themselves. In particular, lack of a persistent memory poses a significant challenge for models strategizing to subvert security.
In a new paper, we introduce evaluations that test strategic capability in the context of particular simulated deployments. We argue that in many deployments, a stateless model would need to reason highly reliably in order for its strategizing to substantially improve its chances of subverting the control protocol. Our results on Claude 3.5 Sonnet (new) illustrate this dynamic.
Imagine you're an AI [...]
---
Outline:
(04:20) Statelessness and reliability
(06:13) Evaluating Claude's strategic ability from within a simulated deployment
(06:54) Summary of results
(08:09) Upfront Auditing
(13:20) Models can subvert upfront auditing without competently strategizing
(14:29) Strategizing O1-style (using DeepSeek-R1-Lite)
(15:05) Untrusted Monitoring
(17:14) What if models aren't stateless?
(18:46) How conservative is our eval?
(20:47) What do our findings imply about risk from schemers?
---
First published:
December 20th, 2024
Source:
https://redwoodresearch.substack.com/p/measuring-whether-ais-can-statelessly
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Fler avsnitt av Redwood Research Blog
Visa alla avsnitt av Redwood Research BlogRedwood Research Blog med Redwood Research finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.
