
The AI Update X: O3
Om avsnittet
In this AI Update Jhave and Scott discuss O3, the new AI model announced by OpenAI, which excels in high-complexity tasks, demonstrating state-of-the-art capabilities in science, structured reasoning, and problem-solving.
References
Brown, N., & Sandholm, T. (2018). Superhuman AI for heads-up no-limit poker: Libratus beats top professionals. https://www.science.org/doi/10.1126/science.aao1733
Chollet, F. (2019).On the measure of intelligence.https://arxiv.org/abs/1911.01547
DeepSeek-AI et al. (2024). DeepSeek-V3 Technical Report. https://arxiv.org/abs/2412.19437
DeepSeek-AI et al. (2025). DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. https://arxiv.org/abs/2501.12948
Epoch AI. (2024). FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning. https://epoch.ai/frontiermath
Meta Fundamental AI Research Diplomacy Team (FAIR) et al. (2022). Human-level play in the game of Diplomacy by combining language models with strategic reasoning. https://www.science.org/doi/10.1126/science.ade9097
Rein, D., et al. (2023). GPQA: A Graduate-Level Google-Proof Q&A Benchmark. https://arxiv.org/abs/2311.12022
Shao, Z., et al. (2024). DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models. https://arxiv.org/abs/2402.03300
Snell, C., Lee, J., Xu, K., & Kumar, A. (2024). Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters. https://arxiv.org/abs/2408.03314
Sutton, R. (2019). The Bitter Lesson. Incomplete Ideas. http://www.incompleteideas.net/IncIdeas/BitterLesson.html
Xin, H., et al. (2024). DeepSeek-Prover-V1.5: Harnessing Proof Assistant Feedback for Reinforcement Learning and Monte-Carlo Tree Searchhttps://arxiv.org/abs/2408.08152
This work was partially supported by the Research Council of Norway Centres of Excellence, project number 332643, Center for Digital Narrative.
Fler avsnitt
Visa alla avsnitt av Off CenterOff Center med Off Center, a Podcast from the Center For Digital Narrative finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.