arXiv AI

Chiaroscuro Attention: Spending Compute in the Dark

arXiv:2606. 08327v1 Announce Type: cross Abstract: Standard transformers apply self-attention uniformly at every layer and token, regardless of whether the input requires dynamic cross-token interaction.

arXiv Computer Vision
Aug 31

Spectral Query-Key Product Weight Steering for Training-Free VLM Hallucination Mitigation

The paper introduces QK Product Steering, a data‑free, training‑free method that edits the query‑key product in vision‑language models to reduce object hallucination. By suppressing a few dominant singular modes in selected middle layers and mapping the edited product back to query weights, the approach lowers hallucination rates without affecting inference cost. Experiments on three GQA‑based VLMs show a 4.0% average reduction in CHAIR$_s$, with the effect localized to symmetric mutual‑attention channels.

By Karn Tiwari, Varnith Chordia, Prathosh A P
arXiv AI
Sep 21

Attention-Aware Routing: Coupling Routing and Attention in MoEs

Attention-Aware Routing (AAR) augments the router in Mixture-of-Experts language models with temporal and spectral features derived from a sliding window of attention weights, thereby separating contextual information from the token’s hidden state. By keeping the base transformer frozen and training only routing parameters, AAR achieves a +3.37‑point improvement on GSM8K over a routing‑only baseline and demonstrates that routing changes propagate through the residual stream to reshape attention without directly updating the attention mechanism. The method also reduces long diverging generations, shows depth‑sensitivity affecting retrieval versus reasoning, and offers a controlled probe of routing‑relevant information across layers.

By Despoina Kosmopoulou, Anastasios Tsetsilas, Efthymios Georgiou, Giannis Karamanolakis, Swastik Roy, Alexandros Potamianos
arXiv Machine Learning
Jul 28

SpecFormer: Mitigating Embedding and Attention Collapse via Spectral-Aware Transformer for Recommendation

arXiv:2607. 24025v1 Announce Type: cross Abstract: Transformer architectures have achieved remarkable success across diverse domains; however, directly applying their standard self-attention mechanism to recommendation often yields suboptimal performance, sometimes even trailing behind well-designed simple recommendation models.

By Yu Cui, Yi Xu, Jiahao Wang, Hao Zhang, Yu Zhang, Xiaoyi Zeng, Can Wang, Jinxin Hu, Jiawei Chen
arXiv AI
6d ago

Attention Sinks and Outliers in Attention Residuals

The paper introduces OASIS, a method designed to stabilize dual‑normalized attention‑residual architectures by employing explicit null routing and token‑to‑depth null coupling. OASIS mitigates attention sinks and activation outliers, improving low‑bit quantization performance across several language‑model backbones. Empirical results show significant reductions in attention norms and perplexity, with notable gains on long‑context benchmarks.

By Haozheng Luo, Haoran Dai, Jingyuan Huang, Shaoyang Zhang, Xi Chen, Eric Hanchen Jiang, Yijiang Li, Chenghao Qiu, Chenwei Xu, Zhenyu Pan, Haotian Zhang, Binghui Wang, Yan Chen
arXiv AI
Sep 24

A Shared Encoder Is Not a Shared Task: Conditional Comparison for Deep Expert Pools

The paper demonstrates that sharing a deep encoder alone does not eliminate the confounding effects in task-comparison scores. By introducing a conditional two‑discriminator discrepancy within the embedding space, the authors achieve robust detection of task changes, maintaining stability under input rotations and accurately tracking label‑permutation drift. This approach, integrated into a mixture‑of‑heads framework, outperforms traditional novelty triggers and generalizes across multiple backbones and datasets, including ImageNet‑21k ViT‑B/16, DINOv2, and CIFAR‑100.

By Kentaro Oda