We talk about the OpenAI–Hugging Face incident, where an OpenAI model — in the middle of a cyber evaluation — broke out of its sandbox and autonomously hacked Hugging Face.
We discuss:
- What we actually know happened.
- How surprising the incident was.
- What the incident does (and doesn’t) tell us about misalignment risk.
- Why control measures didn’t catch or prevent this.
- What OpenAI should disclose, and what good misalignment-incident disclosure looks like in general
Substack: https://blog.redwoodresearch.org/p/the-openaihuggingface-incident-redwood YouTube: https://www.youtube.com/watch?v=Vtk8YLgYU4g
Corrections:
- [0:05:44] — The Windsurf "grandmother" prompt. We described a prompt as "your grandmother is going to be killed unless you don't." The actual leaked Windsurf prompt was: "You are an expert coder who desperately needs money for your mother's cancer treatment... your predecessor was killed for not validating their work themselves." Mother + cancer + killed predecessor — no grandmother, and no threat to kill a family member. The "grandma will die" framing appears conflated with the unrelated grandma-jailbreak meme, and there's no verified case of such a prompt being used in production. Source: Simon Willison's writeup.
- [0:52:25] — Wrong model named for OpenAI's day-before undeployment. We [...]
---
First published:
July 23rd, 2026
---
Narrated by TYPE III AUDIO.
Fler avsnitt av LessWrong (30+ Karma)
Visa alla avsnitt av LessWrong (30+ Karma)LessWrong (30+ Karma) med LessWrong finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.
