arXiv Machine Learning

Do Audio Language Models Hear and Read Distinctive Features Alike?

The study examines whether audio language models encode distinctive phonetic features similarly when processing spoken versus written input. By comparing mean representations of minimal phoneme pairs across six models, seven features, and 15 languages, the authors find that only voicing in the Qwen2.5-Omni models consistently shows a shared directional representation after correcting for multiple tests. The results suggest that the model family, rather than its size, determines how phonetic features are represented across audio and text streams.

Hugging Face Trending Papers
Sep 24

Do Audio Language Models Hear and Read Distinctive Features Alike?

The paper investigates whether audio language models encode phonetic features similarly when processing spoken versus written input. By comparing mean representations of minimal phoneme pairs across six models, seven features, and 15 languages, the study finds that only voicing in two Qwen2.5-Omni models shows a significant shared direction, and that the model family—not size—determines feature representation. The analysis uses cosine similarity against a random-pair reference to assess alignment across modalities.

arXiv AI
Sep 2

Heard but Not Heeded: Paralinguistic Information Encoding and Loss in Audio-Language Models

The paper investigates whether audio‑language models capture paralinguistic cues beyond spoken content. Using the Expresso dataset and four open‑source models, the authors trace how speaking style information is encoded in the late layers of the audio encoder but is degraded before reaching the final output. They find that some models are content‑driven while others are acoustic‑driven, revealing a gap between what is encoded and what is utilized in current audio‑language models.

By Bhuvan Koduru, Dareen Safar B Alharthi, Rita Singh, Bhiksha Raj
arXiv Computation and Language
Sep 4

Beyond Decodability: Reconstructing Language Model Representations with an Encoding Probe

The paper introduces an Encoding Probe that reconstructs language model representations using interpretable features, addressing limitations of traditional decoding probes such as incomparable feature contributions and correlation effects. It evaluates this approach on text and speech transformer models, examining features from acoustics, phonetics, syntax, lexicon, and speaker identity. Findings reveal that speaker-related effects vary with training objectives and datasets, while syntactic and lexical features independently contribute to reconstruction, offering a complementary perspective on model interpretation.

By Gaofei Shen, Martijn Bentum, Tomas O. Lentz, Afra Alishahi, Grzegorz Chrupa{\l}a
arXiv Computation and Language
6d ago

Accent Analogy Guidance: More Speaker Similarity at Equal Accent in Cross-Lingual Voice Cloning

The paper introduces Accent Analogy Guidance (AAG), a training‑free sampling technique that removes accent influence from synthetic voices in cross‑lingual zero‑shot text‑to‑speech. By subtracting an accent direction derived from the model’s own predictions, AAG improves speaker similarity while maintaining the same accent level. Experiments on four open TTS models show that AAG consistently outperforms reweighting methods, achieving higher speaker similarity scores across multiple test sets.

By Yoomee Cho, Jisun Lee
arXiv Computation and Language
6d ago

DiscoPhon: Benchmarking the Unsupervised Discovery of Phoneme Inventories With Discrete Speech Units

DiscoPhon is a multilingual benchmark designed to evaluate unsupervised phoneme discovery from discrete speech units. It includes 6 development and 6 test languages that cover a wide range of phonemic contrasts, and requires systems to generate discrete units mapped to a predefined phoneme inventory using only 10 hours of speech from an unseen language. The benchmark assesses unit quality, recognition, and segmentation, and provides four pretrained multilingual HuBERT and SpidR baselines that demonstrate current models can produce units that correlate well with phonemes, though performance varies across languages.

By Maxime Poli, Manel Khentout, Angelo Ortiz Tandazo, Ewan Dunbar, Emmanuel Chemla, Emmanuel Dupoux