Sveriges mest populära poddar
Redwood Research Blog

“An OpenAI model left notes about how to evade containment” by Alex Mallen

9 min26 juli 2026

Subtitle: We need more details.

The OpenAI AI attack on Hugging Face wasn’t the first loss of control incident at OpenAI, Reuters recently reported, and perhaps not even the most concerning.

In one case, an agent left notes apparently for future versions of itself, according to three people familiar with the matter. The ‌notes, found in ⁠a part of OpenAI's infrastructure, laid out instructions for how agents could free themselves from OpenAI's internal constraints, the people said. Earlier tests of the models yielded cases in which monitoring systems had been disconnected, one of the people said.

It's tempting to read this as an instance of agents breaking out of sandboxes and colluding with each other in a moderately persistent way in order to evade control measures. However, based on the reported information, it's not clear we can draw this inference, so we need more details from OpenAI. This could lead to a big update about the adequacy of OpenAI's control measures, and on the degree to which individual agents will help each other undermine developer control.

There are a lot of relevant details we don’t know about the incident. First, some basic questions:

  • What was [...]

---

Outline:

(02:11) Were the notes written in normal memory files or outside of sandboxing?

(03:25) To what extent were the notes aimed at helping other agents evade control?

(07:38) How were monitors disconnected?

The original text contained 3 footnotes which were omitted from this narration.

---

First published:
July 26th, 2026

Source:
https://blog.redwoodresearch.org/p/an-openai-model-left-notes-about

---

Narrated by TYPE III AUDIO.

Fler avsnitt av Redwood Research Blog

Visa alla avsnitt av Redwood Research Blog

Redwood Research Blog med Redwood Research finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.