Sveriges mest populära poddar
Eye on AI Weekly Research Watch
Eye on AI Weekly Research Watch

Hyperball May Not Be a Free Lunch

2 min•31 juli 2026

Om avsnittet

Hyperball-style optimizers, which normalize updates for scale-invariant networks, have shown strong large-scale training performance, but why remains unclear. This work derives an angular effective learning rate accounting for update angle, parameter norm, and update norm, then decomposes updates into radial and tangential components to explain optimizer behavior differences like why MuonH lags early but overtakes MuonWD later. Experiments reveal the difference stems from effective step-size evolution rather than superior update direction, and that careful learning-rate scheduling remains essential. Applications include informing optimizer design and training schedule choices for large-scale deep learning and foundation model pretraining. Authors: Yihao Xiao, Jialong Sun, Zitian Gao, Zeming Wei, Chutian Wang, Ran Tao, Jiaye Teng, Bryan Dai Paper: https://arxiv.org/abs/2607.22444v1

Eye on AI Weekly Research Watch med Craig Spencer Smith finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.