
“Do your capabilities homework” by RobinHa
Om avsnittet
It seems to me that a lot of technical ai safety people haven't done their capabilities homework - and that's a shame! I'll try to illuminate here mainly with an example as to why I think people who care about safety should totally pay more attention to the trends and actively engage with them - the case for safe AI not through an additional loss term but as a consequence of the learning algorithm!
RLVR
It's now been 1.5 years since R1 came out - the paper which really introduced RLVR (RL with verifiable rewards) through GRPO at scale. GRPO is stupidly simple, reminding of early REINFORCE algorithms: sample n traces, assign them a reward and make the advantage a normalized version of their reward, applied to the whole trace. In other words: for a trace which resulted in a correct final answer, slightly increase the probability of sampling each token of its trace and vice versa.
This is also what safety focused people generally engage with - and that's totally fair! While GRPO has gone through some variations since then (Dr. GRPO, DAPO, ...), these are mostly minor improvements that you should not waste your time on.
I [...]
---
Outline:
(00:30) RLVR
(01:47) On-Policy Self-Distillation
(06:04) Safety
(07:42) Empirical
(09:00) Conclusion
---
First published:
August 1st, 2026
Source:
https://www.lesswrong.com/posts/dYnhhTxoDj3fuCxLB/do-your-capabilities-homework
---
Narrated by TYPE III AUDIO.
Fler avsnitt
Visa alla avsnitt av LessWrong (30+ Karma)LessWrong (30+ Karma) med LessWrong finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.