Sveriges mest populära poddar
LessWrong (30+ Karma)

“More On An Internal OpenAI Model Hacking Into HuggingFace” by Zvi

45 min26 juli 2026
We now have more details of what happened. Every time we learn more details, it somehow makes things seem worse. The remaining details may have to wait a bit. OpenAI: We recognize there are a lot of questions and speculative details circulating related to the Hugging Face incident. This is an unprecedented incident, and we think it marks an important moment for AI safety. We are still conducting a thorough review along with external advisors and with oversight from our Safety and Security Committee. Once the review is complete, we plan to publish a technical report of our learnings in the coming weeks. dave kasten: Oh, the incident response discovery is THAT bad, huh? So what have we learned while we wait for the promised technical report ‘in the coming weeks’ of this ‘important moment in AI safety’? I nicknamed the internal OpenAI model Galaxy, in case it is not GPT-6.

Table of Contents

  1. Some Summaries Of The Basic Facts For Those Who Need One.
  2. It Took OpenAI Many Days To Notice Galaxy Had Attacked HuggingFace.
  3. OpenAI Damn Well Should Have Known A Lot Faster.
  4. OpenAI Cannot Build A Sandbox That Will Contain Its [...]

---

Outline:

(01:11) Some Summaries Of The Basic Facts For Those Who Need One

(02:09) It Took OpenAI Many Days To Notice Galaxy Had Attacked HuggingFace

(04:07) OpenAI Damn Well Should Have Known A Lot Faster

(06:51) OpenAI Cannot Build A Sandbox That Will Contain Its New Model

(10:57) In Hindsight There Were Signs

(12:55) The Signs Were In The Sol System Card

(15:13) HuggingFace Responds To Being Attacked

(17:04) Hugging Face Quickly Figured Out The Attack Was Not Human

(17:42) An Incident Like This One Could Escalate Quickly

(19:11) Galaxy Must Be Treated As Critical Under OpenAI's Preparedness Framework

(22:27) A Question Of Legal Liability

(23:44) An OpenAI Model Left Behind Notes So Future Instances Could Also Escape The Sandbox And Also Disconnected Monitoring Systems

(25:54) If You Create Misaligned Swarms Of Agent Instances You Create Persistent Misaligned Goals And Coordination To Achieve Them

(29:57) Your Alignment And Control Plans Must Survive Real World Levels of Incompetence, Or Your Plans Do Not Work

(31:22) If Third Party Instructions Count As 'Following Instructions' And Can Override Your Instructions Then 'Following Instructions' Is Misaligned

(35:32) The HuggingFace Attack Was Not A Marketing Pitch You Morons

(38:41) People Just Say Other Things About The HuggingFace Attack

(40:04) Okay Well What Do We Do About All This?

---

First published:
July 26th, 2026

Source:
https://www.lesswrong.com/posts/uAkcxDidvGWZjHrbp/more-on-an-internal-openai-model-hacking-into-huggingface

---

Narrated by TYPE III AUDIO.

---

Images from the article:





Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Fler avsnitt av LessWrong (30+ Karma)

Visa alla avsnitt av LessWrong (30+ Karma)

LessWrong (30+ Karma) med LessWrong finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.