arXiv:2606. 02136v1 Announce Type: new Abstract: Neural asymmetric routing models increasingly encode directionality through matrix representations and asymmetry-aware attention.
By Li Liang, Jinbiao Chen, Zizhen Zhang
The paper examines why learned gates in sparse attention models offer little advantage over random gates when jointly trained with the transformer. Through experiments on a 31M-parameter transformer, the authors attribute this to routing absorption, where the model’s representations adapt to the imposed mask, diminishing the benefit of learned routing. They also explore hard masking, stochastic mask training, and the impact of trainable attention layers on gate performance, concluding that freezing the model stabilizes routing targets for effective post‑hoc sparsification.
By Keston Aquino-Michaels
Attention-Aware Routing (AAR) augments the router in Mixture-of-Experts language models with temporal and spectral features derived from a sliding window of attention weights, thereby separating contextual information from the token’s hidden state. By keeping the base transformer frozen and training only routing parameters, AAR achieves a +3.37‑point improvement on GSM8K over a routing‑only baseline and demonstrates that routing changes propagate through the residual stream to reshape attention without directly updating the attention mechanism. The method also reduces long diverging generations, shows depth‑sensitivity affecting retrieval versus reasoning, and offers a controlled probe of routing‑relevant information across layers.
By Despoina Kosmopoulou, Anastasios Tsetsilas, Efthymios Georgiou, Giannis Karamanolakis, Swastik Roy, Alexandros Potamianos
arXiv:2609.08189v1 Announce Type: new
Abstract: Dynamic layer routing reduces the inference cost of Large Language Models (LLMs) by learning to skip layers for individual tokens. Existing methods, ho...
By Hongjin Lin, Wentao Wan, Keze Wang
arXiv:2607. 06601v1 Announce Type: cross Abstract: Conditional computation can decouple language model quality from per-token inference cost, yet leading techniques act on a single axis in isolation: Mixture-of-Experts (MoE) sparsifies the FFN, Mixture-of-Depths (MoD) skips whole transformer blocks, and KV-cache quantization compresses attention memory.
By Andrii Balashov, Olena Ponomarova
arXiv:2609.36062v1 Announce Type: new
Abstract: Long-context sequence models face a fundamental tradeoff: softmax attention uses flexible token-level interactions at quadratic cost, whereas linear at...
By Emile Anand, Abdullah Ateyeh, Archer Wang, Marin Solja\v{c}i\'c