Why do bigger neural networks tend to perform better — and by exactly how much? Scaling laws attempt to answer this, but most existing theory relies on simplified assumptions about infinite width or unlimited data. This work studies how generalization error changes as both model width and dataset size vary simultaneously in a tractable two-layer network, revealing a phase diagram with distinct regimes — including a transition into interpolation — governed by the spectral structure of the target function. While theoretical, these findings have practical implications for deciding how to allocate compute budgets, understanding when more data helps more than more parameters, and designing efficient architectures.
Authors: Julius Girardin, Emanuele Troiani, Yizhou Xu, Vittorio Erba, Florent Krzakala, Lenka Zdeborová
Paper: https://arxiv.org/abs/2606.28242v1
Fler avsnitt av Eye on AI Weekly Research Watch
Visa alla avsnitt av Eye on AI Weekly Research WatchEye on AI Weekly Research Watch med Craig Spencer Smith finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.
