arXiv AI By Kasun Dewage, Marianna Pensky, Suranadi De Silva, T. H. Bandara

Spectral Outliers Reveal Dominant Learned Structure in Transformer Attention

Read the original on arXiv AI →

arXiv:2608. 07921v1 Announce Type: cross Abstract: We apply Marchenko-Pastur (MP) random matrix theory to pre-trained attention weights in order to separate each projection matrix into a random-like bulk and a set of spectral outliers.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
4d ago

Pattern Formation in Transformers

arXiv:2609.37921v1 Announce Type: new Abstract: What are the inductive biases of a Transformer architecture? Existing theory on how the forward pass shapes representations either considers whether Tr...

By Erkan Turan, Gaspard Abel, Maks Ovsjanikov
arXiv Machine Learning
Sep 11

Optimizing Three Critical Factors for Practical and Effective OOD Detection Fine-Tuning

The paper introduces a framework for out-of-distribution (OOD) detection that addresses the trade‑off between detection performance and classification accuracy caused by fine‑tuning with auxiliary outlier data. It optimizes three factors—model reminder, data sampling, and representation learning—by proposing Self‑Knowledge Distillation to preserve accuracy, Semi‑hard Outlier Sampling to enhance detection with minimal data, and Outlier‑aware Supervised Contrastive Learning to improve ID‑OOD separability. The combined approach yields cumulative gains, outperforming existing methods on diverse benchmarks, especially in long‑tailed scenarios, and offers a robust baseline for real‑world OOD detection.

By Hyunjun Choi, JaeHo Chung, Hawook Jeong
arXiv Machine Learning
Sep 23

Magnitude Profile Pruning: Calibration-Free Structured Attention Head Removal for Transformer Compression

Magnitude Profile Pruning introduces a training‑free, calibration‑free method for removing attention heads in Transformer models by statistically detecting outliers in weight row norms. Heads whose projection weights fall within the bulk of the distribution are pruned, while outlier heads are retained. Across several models, the MP‑G variant achieves superior perplexity at various sparsity levels and yields significant parameter and FLOP reductions without requiring forward passes, calibration data, or gradient computations.

By Kasun Dewage, Marianna Pensky, Heranga K. Rathnasekara, Suranadi De Silva