arXiv AI

Bias Analysis of L2 Speaking Assessment Systems Using Concept Activation Vectors

arXiv:2608. 06300v1 Announce Type: new Abstract: Automatic speaking assessment systems are increasingly deployed in high-stakes settings to mark second language (L2) learners' speaking tests, making it critical to show that their scores depend on speaking proficiency rather than irrelevant speaker attributes such as first language (L1) or age.

arXiv Machine Learning
1d ago

Beyond Linear Concepts: Discovering and Aligning Non-Linear Concept Manifolds in Large Language Models

The paper extends mechanistic interpretability of large language models by modeling concepts as low‑dimensional non‑linear manifolds rather than linear subspaces. It introduces a concept‑based alignment (CBA) score to compare these manifolds across layers and models, revealing block structures in intermediate layers, a shift from syntax‑dominated to mixed syntactic‑semantic concepts, and training‑dependent multilingual sharing. The study also shows that alignment patterns differ across model families and training stages, with adjacent stages aligning more closely than distant ones.

By Tido Specht, Elias Benedict Krey, Nils Neukirch, Nils Strodthoff
arXiv Computation and Language
Sep 4

Beyond Decodability: Reconstructing Language Model Representations with an Encoding Probe

The paper introduces an Encoding Probe that reconstructs language model representations using interpretable features, addressing limitations of traditional decoding probes such as incomparable feature contributions and correlation effects. It evaluates this approach on text and speech transformer models, examining features from acoustics, phonetics, syntax, lexicon, and speaker identity. Findings reveal that speaker-related effects vary with training objectives and datasets, while syntactic and lexical features independently contribute to reconstruction, offering a complementary perspective on model interpretation.

By Gaofei Shen, Martijn Bentum, Tomas O. Lentz, Afra Alishahi, Grzegorz Chrupa{\l}a
arXiv Machine Learning
Jun 18

From Sparse Features to Trustworthy Proxies: Certifying SAE-Based Interpretability

arXiv:2606. 18383v1 Announce Type: new Abstract: Sparse autoencoders (SAEs) are increasingly used to extract interpretable features from language models (LMs), yet a central question remains: when can an SAE-based explanation be treated as a faithful view of an underlying frozen LM We study this through a post-hoc generalization framework that certifies the LM via a sparse proxy, obtained by replacing a native hidden activation with its pretrained SAE reconstruction.

By Dibyanayan Bandyopadhyay, Asif Ekbal