arXiv Computation and Language

Evidence for systematic semantic structure in individual letters

arXiv Computation and Language
Aug 31

Tracing the complexity profiles of different linguistic phenomena through the intrinsic dimension of LLM representations

The paper investigates the intrinsic dimension (ID) of large language model (LLM) representations as an indicator of linguistic complexity. By comparing ID across model layers for coordination vs. subordination, right‑branching vs. center‑embedding, and unambiguous vs. ambiguous attachment, the authors find consistent ID differences that align with established complexity contrasts. Experiments across six LLMs, including representational similarity and layer pruning analyses, confirm that more complex phenomena produce higher ID profiles, with peaks occurring at different layers for each contrast.

By Marco Baroni, Emily Cheng, Iria de-Dios-Flores, Francesca Franzon
arXiv AI
Jul 7

The Rise of Verbal Tics in Large Language Models: A Systematic Analysis Across Frontier Models

arXiv:2604. 19139v3 Announce Type: replace-cross Abstract: As Large Language Models (LLMs) continue to evolve through alignment techniques such as Reinforcement Learning from Human Feedback (RLHF) and Constitutional AI, a growing and increasingly conspicuous phenomenon has emerged: the proliferation of verbal tics--repetitive, formulaic linguistic patterns that pervade model outputs.

By Shuai Wu, Xue Li, Yanna Feng, Yufang Li, Zhijun Wang, Ran Wang
arXiv Computation and Language
Sep 14

Quantifying Consonant Contributions to Word Intelligibility via Acoustic Masking

The study introduces a scalable acoustic‑masking method to quantify how much each consonant contributes to word intelligibility. By silencing individual consonants in isolated words and measuring misrecognition rates with three ASR models, the authors define a mask‑induced misrecognition rate (MMR). Across English, Spanish, German, and Czech, MMR negatively correlates with phoneme frequency and positively with functional load, revealing that consonant importance varies by language.

By Eunjung Yeo, Kwanghee Choi, Krupaben Kothadia, Visar Berisha, Julie M. Liss, David R. Mortensen, David Harwath
arXiv AI
Aug 28

How Unlikely Is "Unlikely"? Assessing Verbal Probability Perception Across Large Language Models

The study evaluates how large language models (LLMs) interpret verbal probability expressions by mapping words to numbers and testing consistency across 19 models. Results show that LLMs largely mirror human benchmarks—preserving word order, recovering key anchor points, and reflecting the high variance of the term "possible"—but they exhibit a systematic upward bias for negative expressions like "unlikely" and "improbable." Explanation elicitation reduces within‑model variance but increases divergence between models, while a bidirectional roundtrip test reveals that leading models maintain coherent internal representations.

By Christos Petridis, Konstantinos Pelechrinis, Zoran Obradovic
arXiv Computer Vision
Aug 27

Do Vision-Language Models Agree on the Affective Qualities of Shape? A Cross-Model Audit for Generative Design Interfaces

The study audits six vision‑language models (VLMs) to assess whether they consistently encode affective qualities of 3D shapes, using Kansei adjective pairs as affective axes. Across ten ShapeNet categories, models show moderate agreement (mean rank correlation 0.36) that is lower than geometric controls but higher than unrelated adjective pairs, with convergence varying widely by category and axis. The authors demonstrate how this audit informs a UI prototype that selectively exposes Kansei descriptors for generative design interfaces.

By Luca Bux, Thiago Rios, Ingo Scholtes, Stefan Menzel
Hugging Face Trending Papers
Aug 18

Language Has Two Parameters: Narrative-Induced Semantic Plasticity and Phase-Sensitive Interpretation

The paper argues that language operates with two parameters: amplitude, which measures how often words co‑occur, and phase, a signed relational factor that determines how co‑activated meanings combine and can reverse a meaning’s contribution. Unlike amplitude, phase is not captured by standard word embeddings or transformer attention weights and is indexed to individuals and dyadic interactions. The authors propose six empirical predictions to test phase’s role and suggest that future language models should incorporate agent‑indexed, phase‑bearing semantic states.