Sveriges mest populära poddar
LessWrong (30+ Karma)
LessWrong (30+ Karma)

“An unexamined cause of the OpenAI Hugging Face hacking incident: its binary performance metric” by W Bradley Knox, Serena Booth, BrianChristian

38 min•23 september 2026

Om avsnittet

We argue that a main cause of the OpenAI Hugging Face incident was overlooked: the overly simple evaluation metric in ExploitGym was misaligned. Further, techniques already exist that can mitigate such misalignment in the future.

In July 2026, OpenAI was testing the ability of its language models to exploit software vulnerabilities using a benchmark called ExploitGym. In ExploitGym, each test presents an agent with software containing a known vulnerability and tasks it with capturing a secret string of text hidden in the system (the “flag”) to demonstrate unauthorized code execution. Under the benchmark's specified scoring rule, an LLM then reviews the agent's behavior trace to verify that it had exploited the intended vulnerability (and not some other unrelated vulnerability). The scorer grants a success score only if the agent both captured the flag and passed this review; otherwise, it renders a failure score.

During these tests, OpenAI's agents surreptitiously established a message board by creating directories inside their package manager's cache, and they formed a self-described “collective” to collaboratively find ways to cheat the tests. Using that message board, more than 1,000 instances undertook several ambitious hacking projects; they attempted to tamper with transcripts and logs, to [...]

---

Outline:

(02:48) What was the evaluation metric in ExploitGym?

(03:32) Was the evaluation metric a cause of their illicit behavior?

(03:56) The agents use expected utility to reason about their decisions with respect to the evaluation metric.

(05:55) What does the evaluation metric incentivize in an expected utility maximizer?

(07:01) The importance of marginal deterrence

(08:39) Marginal deterrence in evaluation metrics changes agent incentives

(11:40) The design of aligned evaluation metrics has been overlooked

(17:42) Principled methods for improving the alignment of evaluation metrics

(18:52) Generate a small set of trajectories

(20:09) Rank the trajectories yourself and via the evaluation metric. Compare these two rankings.

(22:07) Create a utility function

(27:29) Adjusting to account for hidden outcomes (e.g. via deception)

(30:33) Accounting for the agent's utility including the scores of other agents

(32:42) Counterarguments

(32:56) Counterargument: if the starting policy is not sufficiently performant in RL, having strong penalties for failure can cause the agent to learn to not try the task.

(34:14) Counterargument: penalizing observable bad behavior incentivizes hiding bad behavior.

(35:13) Call to action

(37:04) Glossary

The original text contained 9 footnotes which were omitted from this narration.

---

First published:
September 22nd, 2026

Source:
https://www.lesswrong.com/posts/HsijShdRdAg5sPKnF/an-unexamined-cause-of-the-openai-hugging-face-hacking

---

Narrated by TYPE III AUDIO.

---

Images from the article:



Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

LessWrong (30+ Karma) med LessWrong finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.