Sveriges mest populära poddar
Redwood Research Blog

“To be legible, evidence of misalignment probably has to be behavioral” by Ryan Greenblatt

6 min15 april 2025

Subtitle: Evidence from just model internals (e.g. interpretability) is unlikely to be broadly convincing..

One key hope for mitigating risk from misalignment is inspecting the AI's behavior, noticing that it did something egregiously bad, converting this into legible evidence the AI is seriously misaligned, and then this triggering some strong and useful response (like spending relatively more resources on safety or undeploying this misaligned AI).

You might hope that (fancy) internals-based techniques (e.g., ELK methods or interpretability) allow us to legibly incriminate misaligned AIs even in cases where the AI hasn't (yet) done any problematic actions despite behavioral red-teaming (where we try to find inputs on which the AI might do something bad), or when the problematic actions the AI does are so subtle and/or complex that humans can't understand how the action is problematic1. That is, you might hope that internals-based methods allow us to legibly incriminate [...]

---

First published:
April 15th, 2025

Source:
https://redwoodresearch.substack.com/p/to-be-legible-evidence-of-misalignment

---

Narrated by TYPE III AUDIO.

Fler avsnitt av Redwood Research Blog

Visa alla avsnitt av Redwood Research Blog

Redwood Research Blog med Redwood Research finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.