arXiv Computation and Language
Aug 31

Tracing the complexity profiles of different linguistic phenomena through the intrinsic dimension of LLM representations

The paper investigates the intrinsic dimension (ID) of large language model (LLM) representations as an indicator of linguistic complexity. By comparing ID across model layers for coordination vs. subordination, right‑branching vs. center‑embedding, and unambiguous vs. ambiguous attachment, the authors find consistent ID differences that align with established complexity contrasts. Experiments across six LLMs, including representational similarity and layer pruning analyses, confirm that more complex phenomena produce higher ID profiles, with peaks occurring at different layers for each contrast.

By Marco Baroni, Emily Cheng, Iria de-Dios-Flores, Francesca Franzon
arXiv AI
Jul 7

The Rise of Verbal Tics in Large Language Models: A Systematic Analysis Across Frontier Models

arXiv:2604. 19139v3 Announce Type: replace-cross Abstract: As Large Language Models (LLMs) continue to evolve through alignment techniques such as Reinforcement Learning from Human Feedback (RLHF) and Constitutional AI, a growing and increasingly conspicuous phenomenon has emerged: the proliferation of verbal tics--repetitive, formulaic linguistic patterns that pervade model outputs.

By Shuai Wu, Xue Li, Yanna Feng, Yufang Li, Zhijun Wang, Ran Wang
arXiv Computation and Language
Sep 14

Quantifying Consonant Contributions to Word Intelligibility via Acoustic Masking

The study introduces a scalable acoustic‑masking method to quantify how much each consonant contributes to word intelligibility. By silencing individual consonants in isolated words and measuring misrecognition rates with three ASR models, the authors define a mask‑induced misrecognition rate (MMR). Across English, Spanish, German, and Czech, MMR negatively correlates with phoneme frequency and positively with functional load, revealing that consonant importance varies by language.

By Eunjung Yeo, Kwanghee Choi, Krupaben Kothadia, Visar Berisha, Julie M. Liss, David R. Mortensen, David Harwath