Sveriges mest populära poddar
LessWrong (30+ Karma)

“AI #178: A Fire Alarm For General Intelligence” by Zvi

1 tim 25 min24 juli 2026

The story that matters most this week is that OpenAI's internally deployed models have severe alignment problems, including repeatedly breaking out of their sandboxes, and in one case sending a swarm of agents that broke into HuggingFace in order to steal the answers to the benchmark ExploitGym.

It is much more important that you read those two posts, and the one on Kimi K3, than to read this one that rounds up the other news of the week.

OpenAI wants to present this as largely an infrastructure and safeguards problem, that it needs to build more secure sandboxes and have better supervision. It does need to do those things, and those are indeed problems, but no that is not the problem.

The problem is severe misalignment, which by default will only get worse.

Our methods of training highly capable LLMs, especially at OpenAI but also everywhere else, lead to systematic misalignment of exactly the type LessWrong has been worried about for a long time. We know some of the causes, and some of the mistakes we need to avoid when doing RL that rewards misaligned behaviors including reward hacking, but we do not know how [...]

---

Outline:

(03:42) Language Models Offer Mundane Utility

(04:24) Language Models Don't Offer Mundane Utility

(07:38) Fable Disproves The Jacobian Conjecture Via Counterexample

(11:24) Claude Fable Will Remain In Max Plan Indefinitely

(13:39) Huh, Upgrades

(14:42) On Your Marks

(19:48) Deepfaketown and Botpocalypse Soon

(20:42) Fun With Media Generation

(20:51) Cyber Lack of Security

(22:07) They Took Our Jobs

(22:56) Get Involved

(24:47) Introducing

(25:46) In Other AI News

(28:02) More on Kimi K3

(33:08) Show Me the Money

(33:55) Quiet Speculations

(37:35) Potential Trouble At UK AISI

(39:29) Pick Up The Phone

(40:30) OpenAI Has Some Alignment Problems

(46:48) The Quest for Sane Regulations

(52:02) Chip City

(53:10) The Week in Audio

(53:27) People Just Say Things

(56:42) Rhetorical Innovation

(58:34) The Rome Declaration

(01:04:02) Aligning a Smarter Than Human Intelligence is Difficult

(01:07:52) Anthropic Surveys Things It Calls Misalignment

(01:13:33) Cooperative Alignment

(01:17:54) Other People Are Not As Worried About AI Killing Everyone

(01:19:35) The Lighter Side

---

First published:
July 23rd, 2026

Source:
https://www.lesswrong.com/posts/BK7E4jHNMykpnt796/ai-178-a-fire-alarm-for-general-intelligence

---

Narrated by TYPE III AUDIO.

---

Images from the article:















Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Fler avsnitt av LessWrong (30+ Karma)

Visa alla avsnitt av LessWrong (30+ Karma)

LessWrong (30+ Karma) med LessWrong finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.