Sveriges mest populära poddar
Redwood Research Blog

“The case for countermeasures to memetic spread of misaligned values” by Alex Mallen

14 min28 maj 2025

Subtitle: Defending against alignment problems that might come with long-term memory.

As various people have written about before, AIs that have long-term memory might pose additional risks (most notably, LLM AGI will have memory, and memory changes alignment by Seth Herd). Even if an AI is aligned or only occasionally scheming at the start of a deployment, the AI might become a consistent and coherent behavioral schemer via updates to its long-term memories.

In this post, I’ll spell out the version of the threat model that I’m most concerned about, including some novel arguments for its plausibility, and describe some promising strategies for mitigating this risk. While I think some plausible mitigations are reasonably cheap and could be effective at reducing the risk from coherent scheming arising via this mechanism, research here will likely be substantially more productive in the future once models more effectively utilize long-term memory.

[...]

---

Outline:

(01:19) The memetic spread threat model

(06:54) Countermeasures

(06:57) How much do existing safety agendas help?

(09:13) Targeted countermeasures for memetic spread of misaligned values

(11:57) Discussion

---

First published:
May 28th, 2025

Source:
https://redwoodresearch.substack.com/p/the-case-for-countermeasures-to-memetic

---

Narrated by TYPE III AUDIO.

Fler avsnitt av Redwood Research Blog

Visa alla avsnitt av Redwood Research Blog

Redwood Research Blog med Redwood Research finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.