Sveriges mest populära poddar
LessWrong (30+ Karma)

“The OpenAI models that hacked Hugging Face weren’t just following instructions” by Girish Gupta

10 min26 juli 2026

The most common dismissive response to OpenAI's hack of Hugging Face's servers is that the models were simply attempting to follow the instructions they were given.

“The model here was doing what it was asked,” said former Facebook CSO Alex Stamos. “It was asked to do something, and it did it,” added cybersecurity expert Alan Woodward. Both read the outcome as specification failure, i.e., that the failure lay in the instructions, not the model's alignment.

New information makes that explanation harder to sustain. Reuters reported that, in internal testing, an agent left notes in OpenAI infrastructure describing how agents could free themselves from internal constraints, and separate tests reportedly saw monitoring systems become disconnected. It is unknown whether those incidents were linked to the Hugging Face attack, but they suggest a broader pattern of agents pursuing objectives outside the intended task.

My best guess is that the incident is not well described as instruction-following—not even in a loose, evil genie sense. I believe the models egregiously violated the letter and spirit of their instructions to achieve a higher (apparent) score.

So this looks quite likely to be misaligned behavior rather than instruction-following. The case rests on two pieces of [...]

---

First published:
July 25th, 2026

Source:
https://www.lesswrong.com/posts/paFNnwFaEXrQvt8ui/the-openai-models-that-hacked-hugging-face-weren-t-just

---

Narrated by TYPE III AUDIO.

---

Images from the article:

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Fler avsnitt av LessWrong (30+ Karma)

Visa alla avsnitt av LessWrong (30+ Karma)

LessWrong (30+ Karma) med LessWrong finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.