
Eye on AI Weekly Research Watch
How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks
2 min•16 september 2026
Om avsnittet
Physics faculty and graduate students re-graded frontier model answers on six physics benchmarks. Most answers marked wrong turned out to reflect bad reference solutions, grading errors or unclear questions. After fixes, GPT-5.6-Sol rose from 47.3% to 78.7% on HLE-Physics and from 61.0% to 87.2% on CMT-Benchmark. The authors say these benchmarks understate what the models can do.
Authors: Ali Ansari, Haoran Sun, Andy Zeyi Liu, Mark Jabbour, Yongshan Ding, et al., Lucas Baker, Arman Cohan, John Sous (51 authors)
Paper: https://arxiv.org/abs/2609.13009v1
Eye on AI Weekly Research Watch med Craig Spencer Smith finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.