arXiv:2609.18533v1 Announce Type: new
Abstract: Automatic speech recognition (ASR) systems exhibit unequal error rates across speaker groups, motivating interventions on their internal representation...
By Nicolas Bourrel, Abderrahmane Issam, Gerasimos Spanakis
Automatic speech recognition (ASR) systems exhibit unequal error rates across speaker groups, motivating interventions on their internal representations. We ask whether speaker-linked attributes that...
The paper extends mechanistic interpretability of large language models by modeling concepts as low‑dimensional non‑linear manifolds rather than linear subspaces. It introduces a concept‑based alignment (CBA) score to compare these manifolds across layers and models, revealing block structures in intermediate layers, a shift from syntax‑dominated to mixed syntactic‑semantic concepts, and training‑dependent multilingual sharing. The study also shows that alignment patterns differ across model families and training stages, with adjacent stages aligning more closely than distant ones.
By Tido Specht, Elias Benedict Krey, Nils Neukirch, Nils Strodthoff
The paper introduces an Encoding Probe that reconstructs language model representations using interpretable features, addressing limitations of traditional decoding probes such as incomparable feature contributions and correlation effects. It evaluates this approach on text and speech transformer models, examining features from acoustics, phonetics, syntax, lexicon, and speaker identity. Findings reveal that speaker-related effects vary with training objectives and datasets, while syntactic and lexical features independently contribute to reconstruction, offering a complementary perspective on model interpretation.
By Gaofei Shen, Martijn Bentum, Tomas O. Lentz, Afra Alishahi, Grzegorz Chrupa{\l}a
arXiv:2608.29034v1 Announce Type: cross
Abstract: A wide range of methods have been proposed for interpreting language models, delivering important insights into their inner workings. However, differ...
By Zhang Enyan, R. Thomas McCoy
arXiv:2606. 18383v1 Announce Type: new Abstract: Sparse autoencoders (SAEs) are increasingly used to extract interpretable features from language models (LMs), yet a central question remains: when can an SAE-based explanation be treated as a faithful view of an underlying frozen LM We study this through a post-hoc generalization framework that certifies the LM via a sparse proxy, obtained by replacing a native hidden activation with its pretrained SAE reconstruction.
By Dibyanayan Bandyopadhyay, Asif Ekbal