arXiv AI

A Controlled Study of Attention-Only Transformers

arXiv:2607. 18363v1 Announce Type: cross Abstract: Feed-forward networks hold two thirds of a transformer's non-embedding parameters, yet the architecture has not received a necessity test that controls parameters, compute, and depth at once.

arXiv AI
Sep 21

Attention-Aware Routing: Coupling Routing and Attention in MoEs

Attention-Aware Routing (AAR) augments the router in Mixture-of-Experts language models with temporal and spectral features derived from a sliding window of attention weights, thereby separating contextual information from the token’s hidden state. By keeping the base transformer frozen and training only routing parameters, AAR achieves a +3.37‑point improvement on GSM8K over a routing‑only baseline and demonstrates that routing changes propagate through the residual stream to reshape attention without directly updating the attention mechanism. The method also reduces long diverging generations, shows depth‑sensitivity affecting retrieval versus reasoning, and offers a controlled probe of routing‑relevant information across layers.

By Despoina Kosmopoulou, Anastasios Tsetsilas, Efthymios Georgiou, Giannis Karamanolakis, Swastik Roy, Alexandros Potamianos
arXiv Computation and Language
3d ago

Learning Functional Subspaces for Neural Network Compression

arXiv:2609.40127v1 Announce Type: cross Abstract: Modern transformers pair impressive capabilities with substantial memory and compute demands. Low-rank weight factorization reduces both while keepin...

By Massimo Bini, Anders Christensen, Stephan Alaniz, Judah Goldfeder, Ole Winther, Yann LeCun, Ravid Shwartz-Ziv, Zeynep Akata
arXiv AI
6d ago

Attention Sinks and Outliers in Attention Residuals

The paper introduces OASIS, a method designed to stabilize dual‑normalized attention‑residual architectures by employing explicit null routing and token‑to‑depth null coupling. OASIS mitigates attention sinks and activation outliers, improving low‑bit quantization performance across several language‑model backbones. Empirical results show significant reductions in attention norms and perplexity, with notable gains on long‑context benchmarks.

By Haozheng Luo, Haoran Dai, Jingyuan Huang, Shaoyang Zhang, Xi Chen, Eric Hanchen Jiang, Yijiang Li, Chenghao Qiu, Chenwei Xu, Zhenyu Pan, Haotian Zhang, Binghui Wang, Yan Chen
arXiv Computation and Language
Sep 1

On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability

arXiv:2608.30320v1 Announce Type: new Abstract: We describe the architecture and ablations of Qwen3.8-Flash-Next, a sparse mixture-of-experts model with 125B parameters, 6B activated per token, and a...

By Zihan Qiu, Zekun Wang, Xiao Li, Yanpeng Li, Yang Xu, Yixuan Wang, Huaqing Zhang, Rui Men, Bochao Mao, Chengruidong Zhang, Fan Zhou, Hao Luo, Haofeng Huang, Haoran Lian, Haoyan Huang, Hongqing Chen, Jianwei Zhang, Jing Xu, Junjie Wang, Langshi Chen, Liangyu Wang, Linlang Jiang, Man Yuan, Minmin Sun, Peng Jin, Siqi Zhang, Siyu Wang, Xingzhang Ren, Yakai Wang, Yi Zhang, Yiming Dong, Yizhong Cao, Yubo Ma, Yunfei Mao, Bo Zheng, Dayiheng Liu