arXiv AI By Viet-Hoang Tran, Vinh Khanh Bui, Van-Hoan Trinh, Tan Lai Ngoc, Tan M. Nguyen

Functional Equivalence in Attention: A Comprehensive Study with Applications to Linear Mode Connectivity

Read the original on arXiv AI →

arXiv:2606. 17830v1 Announce Type: cross Abstract: Neural network parameter spaces are inherently non-injective, as distinct parameter configurations can realize identical functions through functional equivalence.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 4

Modern Transformers Are Implicit Hybrids: From Functional Differentiation to Principled Hybrid Architecture Design

The paper proposes a principled way to design hybrid transformer architectures that combine Full Attention (FA) and Linear Attention (LA). By introducing two intervention metrics—RoPE Frequency Importance Score (RFIS) and RoPE Positional Dependence (RPD)—the authors identify a clear taxonomy of retrieval and positional heads, defining a Global Positional Band (GPBand) that aligns with training-length positional scales. Using these insights, they build a Head‑wise Hybrid Architecture (HwH) that assigns FA to global retrieval and LA to local positional modeling, achieving strong language modeling, improved retrieval, and superior zero‑shot long‑context extrapolation compared to standard Transformers and other hybrids.

By Runlin Shi, Bojian Yin, Guoqi Li
arXiv Machine Learning
4d ago

Pattern Formation in Transformers

arXiv:2609.37921v1 Announce Type: new Abstract: What are the inductive biases of a Transformer architecture? Existing theory on how the forward pass shapes representations either considers whether Tr...

By Erkan Turan, Gaspard Abel, Maks Ovsjanikov