arXiv:2510. 00399v2 Announce Type: replace Abstract: The Mamba model has gained significant attention for its computational advantages over Transformer-based models, while achieving comparable performance across a wide range of language tasks.
By Hongkang Li, Songtao Lu, Xiaodong Cui, Pin-Yu Chen, Meng Wang
arXiv:2510. 04212v4 Announce Type: replace-cross Abstract: The pursuit of computational efficiency has driven the adoption of low-precision formats for training transformer models.
By Haiquan Qiu, Quanming Yao
arXiv:2609.37921v1 Announce Type: new
Abstract: What are the inductive biases of a Transformer architecture? Existing theory on how the forward pass shapes representations either considers whether Tr...
By Erkan Turan, Gaspard Abel, Maks Ovsjanikov
The paper introduces a framework for out-of-distribution (OOD) detection that addresses the trade‑off between detection performance and classification accuracy caused by fine‑tuning with auxiliary outlier data. It optimizes three factors—model reminder, data sampling, and representation learning—by proposing Self‑Knowledge Distillation to preserve accuracy, Semi‑hard Outlier Sampling to enhance detection with minimal data, and Outlier‑aware Supervised Contrastive Learning to improve ID‑OOD separability. The combined approach yields cumulative gains, outperforming existing methods on diverse benchmarks, especially in long‑tailed scenarios, and offers a robust baseline for real‑world OOD detection.
By Hyunjun Choi, JaeHo Chung, Hawook Jeong
arXiv:2609.36448v1 Announce Type: new
Abstract: Transformers have demonstrated remarkable in-context learning (ICL) capabilities, enabling them to perform new tasks without additional fine-tuning. Ho...
By Junze Deng, Daouda Sow, Sen Lin, Yingbin Liang
Magnitude Profile Pruning introduces a training‑free, calibration‑free method for removing attention heads in Transformer models by statistically detecting outliers in weight row norms. Heads whose projection weights fall within the bulk of the distribution are pruned, while outlier heads are retained. Across several models, the MP‑G variant achieves superior perplexity at various sparsity levels and yields significant parameter and FLOP reductions without requiring forward passes, calibration data, or gradient computations.
By Kasun Dewage, Marianna Pensky, Heranga K. Rathnasekara, Suranadi De Silva