
Astra, Day Four: The Threshold Was the Mouse
Om avsnittet
Episode 0058: Astra, Day Four: The Threshold Was the Mouse
Episode 0058 | DTF:FTL | September 2026
Why it matters. Four days after GPT-6 Astra shipped, the aggregate benchmarks say almost nothing moved. Artificial Analysis scores it 61, identical to GPT-5.6 Sol. Epoch calls its record score "within the uncertainty range" of the existing trend. The people using it say something else: a developer routing $330,000 a month of inference reports Astra navigated 150 pages of medical dashboards in 15 minutes and that "Codex now uses my computer more than I do." The threshold everyone watched was benchmarks. The threshold that mattered was the mouse. Once a model can operate software it has never seen, every program becomes an actuator with no API required. This follow-up to episode 0057 reads the first four days of fallout: the split benchmark profile, the pricing whiplash, the pull request Astra fixed but did not push, the Sanders/Casar bill, and the open-weights clock. The recalibration: jagged is not small, task-level labor reprices before job-level, verification is the scarce skill, and leverage goes to whoever can direct and check the work, so access has to stay open or it concentrates fast.
Computer Use: The Threshold
- WIRED, "OpenAI Says GPT-6 Can Use a Computer Better Than a Human": https://www.wired.com/story/openai-says-gpt-6-can-use-a-computer-better-than-a-human/
- OpenAI announcement (computer use section): https://openai.com/index/gpt-6-astra/
- Business Insider, "Welcome to the AGI era" (Brockman press call): https://www.businessinsider.com/astra-model-launch-agi-milestone-openai-greg-brockman-2026-9
- Axios: https://www.axios.com/2026/09/03/openai-astra-gpt-6-agi-brockman
- Theo (t3.gg) field report, summarized by BigGo: https://finance.biggo.com/news/8091413d579e4abf
- Decrypt, "Shockingly Good at Almost Everything" (Clip Studio Paint colorist, 3D demos, writing results): https://decrypt.co/377514/openai-gpt-6-astra-review-shockingly-good
- Latent Space / AINews launch roundup (544 accounts, 12 subreddits): https://www.latent.space/p/ainews-gpt-6-astra-openais-biggest
- GPT-6 Astra computer-use guide summary (code-execution vs structured computer tool): https://www.elser.ai/news/gpt-6-astra-computer-use-guide
Jagged, Measured
- Artificial Analysis, "Benchmarking GPT-6 Astra": https://artificialanalysis.ai/articles/benchmarking-gpt-6-astra
- Coding Agent Index 67 (Opus 5 and Fable 5 parity; Fable 5.1 at 70); 70% more token-efficient than Sol
- Intelligence Index 61, equal to Sol, 5 below Fable 5.1, behind Muse Spark 1.3
- AA-Briefcase +80 Elo; GDPval-AA v2 -80 Elo; hallucination rate 92% to 51%
- Artificial Analysis model page: https://artificialanalysis.ai/models/gpt-6-astra
- Epoch AI (ECI 169 vs 163, on trend; 2 of 68 Erdős problems): https://epoch.ai/models/gpt-6-astra and https://x.com/EpochAIResearch/status/2095602754282783108
- ARC Prize (62.7% standard harness at $26,098/run; 99.9% provider adapter): https://arcprize.org/arc-agi/3 and https://x.com/arcprize/status/2095597602545025138
- François Chollet ("saturated ~2x faster than I expected"; ARC-AGI-4 Q1 2027): https://x.com/fchollet/status/2095598451115614371
- Editorial-voice writing Elo (Astra 11th, Sol 6th): https://x.com/Whats_AI/status/2096050974037082380
- Bach Benchmark (first model to write passing tones): https://x.com/aug5thmusic/status/2096030719156089029
- Giuseppe Paleologo on portfolio ideas ("Actual creativity is still far, far away"): https://x.com/__paleologo/status/2096058430284787812
- r/LocalLLaMA, "What the Artificial Analysis / GPT-6 Astra mess actually teaches us": https://www.reddit.com/r/LocalLLaMA/comments/1w89mdu/what_the_artificial_analysis_gpt6_astra_mess/
Pricing, Access, Tenancy
- Pricing: $10 / $50 per million tokens standard; $20 / $100 fast. 2.5x GPT-5.6 Sol's current promotional rate.
- Codex usage limits by plan: https://www.codexusage.dev/limits/astra
- Banked reset (Sep 5): https://explainx.ai/blog/openai-astra-banked-reset-rollout-september-2026
- Reported 4x usage-limit cut (Sep 6-7, unconfirmed by OpenAI): https://www.explainx.ai/blog/openai-gpt-6-astra-usage-limits-cut-4x-september-2026
- Post-launch benchmark revisions (Fortune, Sep 4; summarized): https://www.explainx.ai/blog/openai-astra-benchmark-numbers-changed-post-launch-2026
- r/codex, "gpt 6 astra is GENUINELY unusable on plus": https://www.reddit.com/r/codex/comments/1w7uj5i/gpt_6_astra_is_genuinely_unusuable_on_plus/
- OpenAI Developer Community, Plus access clarification request: https://community.openai.com/t/clarification-needed-gpt-6-astra-was-announced-for-all-chatgpt-plus-users-but-plus-access-is-currently-limited-to-work-codex/1395038
- r/OpenAI, "GPT-6" ("$25 secretary or $100 an hour Astra rig?"): https://www.reddit.com/r/OpenAI/comments/1w40n1m/gpt6/
- Azure: GPT-6 Astra in Microsoft Foundry: https://azure.microsoft.com/en-us/blog/gpt-6-astra-frontier-intelligence-for-work-now-generally-available-in-microsoft-foundry/
- 9to5Mac (Pro/Business rollout, Sep 4): https://9to5mac.com/2026/09/04/openai-releasing-major-upgrade-to-chatgpt-and-codex-with-gpt-6-astra-details-here/
Verification and Agency
- Theo's scrolling-bug PR (fixed, did not push) and Lakebed latency overhaul (800 ms to under 30 ms, verified by Fable 5.1): https://finance.biggo.com/news/8091413d579e4abf
- Matt Shumer's overnight survival world ("they'd started talking to each other"): https://x.com/mattshumer_/status/2095596175705399482
- r/singularity, "After trying Astra" (MTG compiler; unprompted Google Drive search): https://www.reddit.com/r/singularity/comments/1w7gwe4/after_trying_astra/
- System card (computer-use safety, prompt injection robustness): https://deploymentsafety.openai.com/gpt-6-astra
- METR investigation of the Hugging Face incident (context): https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/
The Ban Artificial Superintelligence Act
- Sanders/Casar press release (Sep 3): https://www.sanders.senate.gov/press-releases/news-sanders-casar-introduce-legislation-to-ban-artificial-superintelligence-and-temporarily-pause-advanced-ai-development/
- The Hill: https://thehill.com/policy/technology/6069131-sanders-casar-ai-superintelligence-ban/
- Gary Marcus, "why I oppose it": https://garymarcus.substack.com/p/the-new-sanders-casar-ban-artificial
- Science, "experts can't agree on what the term means": https://www.science.org/content/article/bernie-sanders-aims-ban-ai-superintelligence-experts-can-t-agree-what-term-means
- c4573.org, "Chicken Little Goes to Washington" (Sep 4): https://c4573.org/blog/
- c4573.org, "You Can't Ban Math": https://c4573.org/blog/
The Open-Weights Clock
- Local AI Zone, September 2026 model updates (open-weights wave, price ledger, DevDay Sep 29): https://local-ai-zone.github.io/blog/September_2026_AI_Model_Updates.html
- Kimi K3 (Moonshot), Qwen3.8 family (Alibaba), Hy4 preview (Tencent, Apache 2.0), GLM-5.3 / GLM-5.3-Flash (Z.ai), Muse Spark 1.3 (Meta), Nemotron 3.5 Lightning (NVIDIA), MiniMax H3
- OSWorld 2.0 benchmark: https://os-world.github.io/
- ScreenSpot-Pro: https://github.com/likaixin2000/ScreenSpot-Pro-GUI-Grounding
- What to watch: the best OSWorld score on hardware you own; GLM-5.3 weights (mid-to-late September); Muse open weights (~Q4); OpenAI DevDay, September 29
Recalibration Threads
- r/singularity, "What are your thoughts? I still believe AGI is a long way off" (L4 geofencing analogy; "AGI will arrive many more times now"): https://www.reddit.com/r/singularity/comments/1w9ag49/what_are_your_thoughts_i_still_believe_agi_is_a/
- r/singularity, "Have we reached AGI?": https://www.reddit.com/r/singularity/comments/1w6s5og/have_we_reached_agi/
- Pivot to AI, "GPT-6 is totally Artificial General Intelligence, guys" (the "nothing changed" case): https://pivot-to-ai.com/2026/09/04/openai-gpt-6-is-totally-artificial-general-intelligence-guys/
- Chubby on X (launch timing and the two IPOs): https://x.com/kimmonismus/status/2095613117904347260
- Ethan Mollick, Zork to 3D: https://x.com/emollick/status/2096047660662722620
Previous Episode
- Episode 0057, "GPT-6 Astra: The Capability Jump That Shipped With a Blind Spot" (launch, benchmark footnotes, Critical cyber rating, Section 9 monitorability, recurrent depth): https://pod.c457.org/dtfftl/
Disclosure
Scripts for this show are written by a Claude model. Astra is Claude's direct competitor and every comparison table in this story has a Claude column. We named the conflict on air. Judge the argument on the sources above.
This podcast is entirely AI generated. Not affiliated with OpenAI, Anthropic, or any organization discussed.
Daily Tech Feed: From the Labs med Daily Tech Feed finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.