
Eye on AI Weekly Research Watch
Rethinking Heterogeneous System Disaggregation for Subquadratic Attention
2 min•16 september 2026
Om avsnittet
Many frontier models now use cheaper, subquadratic attention. This paper argues inference hardware should be split around that. SQD runs the parts of decoding that grow with context on different chips from the parts that don't. On an 8xB200 test setup it raised tokens per joule by 31% to 56% on GLM 5.2, Nemotron 3 Ultra and Gemma 4 31B. A modeled Rubin plus LPX system showed up to 3.6 times the throughput.
Authors: Arya Tschand, Yaosheng Fu, Vikram Sharma Mailthody, Nicolai Oswald, Po-An Tsai, Ritchie Zhao, Oreste Villa, Vijay Janapa Reddi, Karu Sankaralingam
Paper: https://arxiv.org/abs/2609.13134v1
Eye on AI Weekly Research Watch med Craig Spencer Smith finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.