Sveriges mest populära poddar
LessWrong (30+ Karma)
LessWrong (30+ Karma)

“Corrigibility Prizes for Existing Work” by Max Harms

14 min•30 september 2026

Om avsnittet

One of my goals for the Corrigibility Research Fund is to retroactively encourage high-quality research on AI alignment (and corrigibility in particular) by awarding prizes. Back in July, I got my feet wet as a fund manager by handing out $27,000 to reward existing work and build interest in the fund. Now, I'd like to disburse an additional $48,000 and use the opportunity to publicly highlight and celebrate the work of the prizewinners from both rounds: about two dozen researchers scattered across roughly a dozen teams.

If the fund continues to be supported in future years, my hope is for prizes like these to become regular, predictable, and large, such that many researchers, year after year, are motivated to aim for them. The awards that I'm announcing here are more ad-hoc than I'd like, and represent only my single perspective trying to balance a wide range of desiderata. Don't take the specific size of each prize purse too seriously. It's all high-quality work. If anyone has ideas for how to improve the retroactive funding process for this kind of scientific work, please leave a comment!

(And as always, if you know of work that I should be aware of [...]

---

Outline:

(02:42) Corrigibility Transformation: Constructing Goals That Accept Updates

(02:49) Rubi Hudson -- $14,000

(04:12) Eval Cooperativeness May Be a Scalable Mitigation for Eval Gaming

(04:18) Jasmine Li and Alex Turner -- $9,000

(05:28) Empowerment, corrigibility, etc. are simple abstractions (of a messed-up ontology)

(05:36) Steven Byrnes -- $6,000

(06:21) Towards Shutdownable Agents: Generalizing Stochastic Choice in RL Agents and LLMs

(06:30) Carissa Cullen, Harry Garland, Alexander Roman, Louis Thomson, Christos Ziakas, Elliott Thornley -- $6,000

(07:42) Assistance with CAST

(07:46) Nathan Helm-Burger -- $6,000

(08:20) The Consciousness Cluster: Emergent preferences of Models that Claim to be Conscious

(08:27) James Chua, Jan Betley, Samuel Marks, Owain Evans -- $6,000

(09:20) ROGUE: Misaligned Agent Behavior Arising from Ordinary Computer Use

(09:27) Jeremy Tien, Abishek Anand, Yu-Rou Tuan, Yuchen Shen, J. Zico Kolter, Aran Nayebi -- $6,000

(10:12) CAST Constitution, Empirical Work on Aspects of Corrigibility that are Unintuitive to LLMs, and other Preliminary Results (Unpublished)

(10:22) Ian Kahn -- $6,000

(10:56) Various Essays on Obedience

(11:00) Seth Herd -- $3,000

(11:49) The corrigibility basin of attraction is a misleading gloss

(11:54) Jeremy Gillen -- $2,000

(12:32) A Structural Similarity Between Two Open Corrigibility Questions and Why Should Corrigible Agents Favor the Present?

(12:40) Ben Saudek -- $2,000

The original text contained 3 footnotes which were omitted from this narration.

---

First published:
September 30th, 2026

Source:
https://www.lesswrong.com/posts/3uJqhrC2idf4eNj5h/corrigibility-prizes-for-existing-work

---

Narrated by TYPE III AUDIO.

---

Images from the article:

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

LessWrong (30+ Karma) med LessWrong finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.