Sveriges mest populära poddar
LessWrong (30+ Karma)
LessWrong (30+ Karma)

[Linkpost] “Estimating GPT-6 Astra’s no-CoT Time Horizon” by Francis Rhys Ward, Dewi Gould

20 min•9 september 2026

Om avsnittet

This is a link post.

TL;DR: We run astra on the task-suite from Think Fast.

  • GPT-6 Astra is close to saturation on our suite, meaning that giving a confident estimate of time horizons (TH) is more challenging.
  • Our rough estimate is that Astra's 50% TH is in [8mins, 1 hour] and probably around 15-40 mins. This is inline with UKAISI's estimate of 30 minutes measured only on math.  
  • Using data from 2019 to April 2026, in Think Fast our median prediction was that no-CoT THs could exceed 7 minutes by 2028. We estimated 30mins by the end of the decade.
  • Astra clearly gets much higher performance on tasks that require serial reasoning, e.g.,
    • Arc-agi-1 and 2, hash, n-hop-look-up, causal-reasoning, sally-anne, all the puzzle tasks.
  • There are limitations with our task suite: Ideally we would have more time-variation (especially longer times) in each specific benchmark, and more benchmarks with longer human completion times.


We use the single-forward pass (31) and generation tasks (6) from the Think Fast suite, totalling 37 benchmarks. Notably we find that Astra gets >= 98% raw accuracy on 10 benchmarks (cf. GPT-5.5 saturates 4 benchmarks – see Appendix). As a result, the [...]

---

Outline:

(04:21) Commentary

(05:22) Acknowledgements

(05:34) Appendix

(05:37) Per-benchmark time horizons

(06:06) Sensitivity to fake benchmarks ablations

(08:23) Raw benchmark performance

(09:06) Table of TH results

(10:28) Implementation caveats

(11:23) GPT-6 Astra: per-benchmark 50% no-CoT time horizons (no filtering or hypothetical benchmarks)

(13:16) GPT-5.5 vs GPT-6 Astra: no-CoT raw accuracy per benchmark

(14:02) Math & science

(14:21) Abstract reasoning

(14:42) Puzzles

(15:02) Language & strategy

(15:21) SWE & cyber

(15:41) Steganography

(16:01) Sabotage & monitoring (TPR @1.5% FPR)

(16:55) Including generation tasks

(17:09) Reasoning token anchor

(17:24) N-hop Task

(17:44) Filler tokens

---

First published:
September 9th, 2026

Source:
https://www.lesswrong.com/posts/ntKx9YHWCwxSeGbRB/estimating-gpt-6-astra-s-no-cot-time-horizon

Linkpost URL:
https://docs.google.com/document/d/1LsJh6ecdfONCSJVpbC44pz1JzcEs4_TM5HD0oihe_70/edit?tab=t.0

---

Narrated by TYPE III AUDIO.

---

Images from the article:











Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

LessWrong (30+ Karma) med LessWrong finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.