
LessWrong (30+ Karma)
“Why research personas despite RL scaling?” by Cleo Nardo
3 min•27 september 2026
Om avsnittet
Here's my rough impression of why people are researching personas despite RL seeming to shape much of the motivations and behaviour of the agents, c.f. Thoughts on the persona selection model (Sam Marks, 24th Sep 2026).
I haven't bothered to check this with anyone.
- Anthropic: "We'll give Claude an aligned persona and hope massive RL doesn't completely burn through it."
- Owain Evans / TruthfulAI: "We'll study personas as part of a broader project of uncovering phenomena in LLM generalisation, which will probably prove useful."
- Center on Long-Term Risk: "Personas may not be enough to build an aligned agent, because RL may play a bigger role in shaping motivations. But personas might be enough to avoid building an anti-aligned agent, i.e. one that is actively malevolent or spiteful."
- David Africa / Resolution (v1): "Scalable oversight protocols like debate may have multiple fixed points, unlike current RL methods which seem more convergent. So the agents' starting dispositions matter. For example, debate might reach a better fixed point, and do so more sample-efficiently, if the agents start honest."
- David Africa / Resolution (v2): "Maybe personas have a 1000-dimensional substructure. If so, we could identify the aligned persona with O(1000) well-chosen datapoints [...]
---
First published:
September 27th, 2026
Source:
https://www.lesswrong.com/posts/i4yswYDSrPFHWpCbi/why-research-personas-despite-rl-scaling
---
Narrated by TYPE III AUDIO.
Fler avsnitt
Visa alla avsnitt av LessWrong (30+ Karma)LessWrong (30+ Karma) med LessWrong finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.