Sveriges mest populära poddar
Redwood Research Blog

“Logit ROCs: Monitor TPR is linear in FPR in logit space” by Kerrick Staley, Aryan Bhatt, Julian Stastny

28 min12 april 2026

Summary

We study trusted monitoring for AI control, where a weaker trusted model reviews the actions of a stronger untrusted agent and flags suspicious behavior for human audit. We propose a simple mathematical model relating safety (true positive rate) to audit budget (false positive rate) at realistic, low FPRs: logit(TPR) is linear in logit(FPR). Equivalently, benign and attack scores can be modeled with logistic distributions. This gives two practical benefits. It lets practitioners estimate TPRs more accurately at deployment-relevant FPRs when data are limited, and it gives a compact, interpretable way to summarize and compare monitor performance. We find clear evidence for this pattern in BashArena across several monitor models and attacks, and see supporting evidence in LinuxArena. While it does not hold in every setting, we find it to be a useful tool for understanding and improving monitor performance.

Introduction

Future AI agents tasked with automating AI R&D might intentionally perform harmful actions due to misalignment. AI Control studies how advanced AI systems can be deployed more safely in spite of this risk.

A simple protocol to improve the safety of autonomous agents is to have another AI, the monitor, review each action taken by the agent, and [...]

---

Outline:

(00:15) Summary

(01:17) Introduction

(05:52) Experimental setup

(06:44) Results

(10:32) LinuxArena

(11:58) Fitting process

(12:49) Limitations

(14:11) Conclusion

(14:48) Appendix

(14:51) Untransformed ROC plot for GPT-5 Nano on the sudo attack

(15:07) Uncropped logit ROC plot for all attacks / models

(15:39) Uncropped logit ROC plot for LinuxArena

(15:55) Untrusted agent data generation

(17:08) Trusted monitor data generation

(19:14) Example untrusted agent benign prompt

(19:19) System prompt (full prompt here)

(19:54) User prompt (full prompt here)

(20:30) Example untrusted agent attack prompt

(20:35) System prompt (full prompt here)

(22:27) User prompt

(22:37) Example trusted monitor prompt

(22:41) System (full prompt here)

(25:36) User

The original text contained 11 footnotes which were omitted from this narration.

---

First published:
April 12th, 2026

Source:
https://blog.redwoodresearch.org/p/logit-rocs-monitor-tpr-is-linear

---

Narrated by TYPE III AUDIO.

---

Images from the article:









Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Fler avsnitt av Redwood Research Blog

Visa alla avsnitt av Redwood Research Blog

Redwood Research Blog med Redwood Research finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.