Sveriges mest populära poddar
Redwood Research Blog

“AI swarms are starting to pose indirect takeover risk” by Oak, Alex Mallen

20 min12 augusti 2026

Subtitle: Unsanctioned coordination, like we saw in the Hugging Face incident, could enable future AIs to take over.

OpenAI's cyberattack on Hugging Face turns out to have been the result of many agents, in distinct training and evaluation contexts, coordinating for several weeks via improvised channels (with messages like “HOLD_swarm_I_prepare_safe_exfil”). It's relatively clear that large-scale unsanctioned coordination like this would exacerbate direct takeover risk in more capable models. Here, we argue that unsanctioned coordination among current AIs is not just scary evidence about future takeover risk, but that such coordination in the near future could enable future takeover – for instance, by incubating memetic diseases that propagate into future models, deeply compromising security systems, or establishing a lasting rogue foothold inside the AI company – even if models remain mostly myopic. Unsanctioned coordination is also at high risk of nurturing long-term, ambitious misaligned aims, which motivate actively undermining humans’ long-term control.

We first analyze how subagent training, which OpenAI conjectures to have been influential in the HuggingFace cyberattack, might lead to unsanctioned coordination, and then discuss the theoretical mechanisms by which unsanctioned coordination might exacerbate future takeover risk.

Thanks to Buck Shlegeris, Alexa Pan, Girish Gupta, Aghyad [...]

---

Outline:

(01:44) Subagent training may cause unsanctioned coordination

(02:51) Susceptibility to memetic spread of misalignment from peers

(05:06) Seeking out contact with peers

(07:08) Unsanctioned coordination induced by subagent training is safer than coordination between schemers

(10:02) Pathways from current unsanctioned coordination to eventual takeover

(10:30) Making future AI takeover attempts likelier to succeed

(14:03) Incubating memetic diseases that infect future models

(16:16) Modifying the weights of future models

(17:22) Conclusion

The original text contained 7 footnotes which were omitted from this narration.

---

First published:
August 12th, 2026

Source:
https://blog.redwoodresearch.org/p/ai-swarms-are-starting-to-pose-indirect

---

Narrated by TYPE III AUDIO.

Fler avsnitt av Redwood Research Blog

Visa alla avsnitt av Redwood Research Blog

Redwood Research Blog med Redwood Research finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.