Sveriges mest populära poddar
Redwood Research Blog

“Behavioral red-teaming is unlikely to produce clear, strong evidence that models aren’t scheming” by Buck Shlegeris

24 min10 oktober 2024

One strategy for mitigating risk from schemers (that is, egregiously misaligned models that intentionally try to subvert your safety measures) is behavioral red-teaming (BRT). The basic version of this strategy is something like: Before you deploy your model, but after you train it, you search really hard for inputs on which the model takes actions that are very bad. Then you look at the scariest model actions you found, and if these contain examples that are strong evidence that the model is scheming (or has other behavioral problems), you don't deploy it.

BRT is likely to be a useful component of a good overall plan for handling risk from scheming. But though I think it's plausible that BRT is able to give us strong evidence that our models are scheming, I don’t think it's going to be able to give us very strong evidence that they aren’t.

A big [...]

---

Outline:

(04:25) A basic version of BRT for schemer risk mitigation

(07:34) BRT with human- or model-generated synthetic inputs

(16:04) But what if you search harder?

(19:22) The adversarial epistemic environment

(20:41) Overall thoughts

---

First published:
October 10th, 2024

Source:
https://redwoodresearch.substack.com/p/behavioral-red-teaming-is-unlikely

---

Narrated by TYPE III AUDIO.

Fler avsnitt av Redwood Research Blog

Visa alla avsnitt av Redwood Research Blog

Redwood Research Blog med Redwood Research finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.