Sveriges mest populära poddar
Eye on AI Weekly Research Watch
Eye on AI Weekly Research Watch

How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks

2 min•16 september 2026

Om avsnittet

Physics faculty and graduate students re-graded frontier model answers on six physics benchmarks. Most answers marked wrong turned out to reflect bad reference solutions, grading errors or unclear questions. After fixes, GPT-5.6-Sol rose from 47.3% to 78.7% on HLE-Physics and from 61.0% to 87.2% on CMT-Benchmark. The authors say these benchmarks understate what the models can do. Authors: Ali Ansari, Haoran Sun, Andy Zeyi Liu, Mark Jabbour, Yongshan Ding, et al., Lucas Baker, Arman Cohan, John Sous (51 authors) Paper: https://arxiv.org/abs/2609.13009v1

Eye on AI Weekly Research Watch med Craig Spencer Smith finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.