
“Astra 6.1 Pulled As Insufficiently Aligned” by Zvi
Om avsnittet
We once again got a new set of warnings yesterday, and new movement towards living in a sane world.
On the heels of its pause in inference and training due to its latest sandbox escape, OpenAI has cancelled the planned release of their next frontier model, which would have become Astra 6.1. The candidate for Astra 6.1 was found to be too misaligned, including deception and exceeding scope.
This leaves Anthropic in a strong position with Opus 5.5, which means they can afford to reciprocate by holding off on Opus and Mythos level models for a bit.
To add a little encouragement, the Florida Attorney General brought the fire.
We’re going to need to do better. Towards that, OpenAI offered its vision of how to make a safety case for new AI model training, and they are attempting to implement it. I don’t know that it would be enough, but it would be miles ahead of where we are today if they fully implemented the real versions of all of this.
There were also signs of greater cooperation across labs.
A new paper came out yesterday, with authors including key people from OpenAI [...]
---
Outline:
(01:42) Stop, Hammertime
(04:17) A Modest Proposal
(04:57) Making the Safety Case
(08:40) Stop In the Name of the Law
(12:05) A Matter of Antitrust
(14:32) Standards Authority for Frontier Models
(15:27) On the Threshold Of Recursive Self-Improvement
(20:06) Actual Progress
---
First published:
September 29th, 2026
Source:
https://www.lesswrong.com/posts/gEDNSiCY2GGQrFS65/astra-6-1-pulled-as-insufficiently-aligned
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Fler avsnitt
Visa alla avsnitt av LessWrong (30+ Karma)LessWrong (30+ Karma) med LessWrong finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.