The paper investigates whether audio language models encode phonetic features similarly when processing spoken versus written input. By comparing mean representations of minimal phoneme pairs across six models, seven features, and 15 languages, the study finds that only voicing in two Qwen2.5-Omni models shows a significant shared direction, and that the model family—not size—determines feature representation. The analysis uses cosine similarity against a random-pair reference to assess alignment across modalities.
The paper investigates whether audio‑language models capture paralinguistic cues beyond spoken content. Using the Expresso dataset and four open‑source models, the authors trace how speaking style information is encoded in the late layers of the audio encoder but is degraded before reaching the final output. They find that some models are content‑driven while others are acoustic‑driven, revealing a gap between what is encoded and what is utilized in current audio‑language models.
By Bhuvan Koduru, Dareen Safar B Alharthi, Rita Singh, Bhiksha Raj
arXiv:2603. 09714v2 Announce Type: replace-cross Abstract: While multi-audio understanding is critical for large audio-language models (LALMs), it remains underexplored.
By Chih-Kai Yang, Yun-Shao Tsai, Yu-Kai Guo, Ping-Le Tsai, Yen-Ting Piao, Hung-Wei Chen, Ting-Lin Hsiao, Yun-Man Hsu, Ke-Han Lu, Hung-yi Lee
The paper introduces an Encoding Probe that reconstructs language model representations using interpretable features, addressing limitations of traditional decoding probes such as incomparable feature contributions and correlation effects. It evaluates this approach on text and speech transformer models, examining features from acoustics, phonetics, syntax, lexicon, and speaker identity. Findings reveal that speaker-related effects vary with training objectives and datasets, while syntactic and lexical features independently contribute to reconstruction, offering a complementary perspective on model interpretation.
By Gaofei Shen, Martijn Bentum, Tomas O. Lentz, Afra Alishahi, Grzegorz Chrupa{\l}a
arXiv:2603. 28378v2 Announce Type: replace-cross Abstract: We present the first systematic Membership Inference Attack (MIA) evaluation of LALMs.
By Jia-Kai Dong, Yu-Xiang Lin, Hung-Yi Lee
arXiv:2609.30483v1 Announce Type: cross
Abstract: Audio language models state numbers for acoustic quantities, and neither human opinion nor a judge model says whether such a number is true of the si...
By Sheng-Tse Lin, Siyuan Zhai, Chien-Liang Kuo, Massa Baali, Bhiksha Raj
arXiv:2608.30853v1 Announce Type: cross
Abstract: While automatic speech recognition (ASR) models have achieved remarkable improvements in recent years, performance disparities persist across differe...
By Ting-Hui Cheng, Line Katrine Harder Clemmensen, Sneha Das
The paper introduces Accent Analogy Guidance (AAG), a training‑free sampling technique that removes accent influence from synthetic voices in cross‑lingual zero‑shot text‑to‑speech. By subtracting an accent direction derived from the model’s own predictions, AAG improves speaker similarity while maintaining the same accent level. Experiments on four open TTS models show that AAG consistently outperforms reweighting methods, achieving higher speaker similarity scores across multiple test sets.
By Yoomee Cho, Jisun Lee
arXiv:2510.00628v3 Announce Type: replace-cross
Abstract: Large audio-language models (LALMs) are often used in tasks that involve reasoning over ordered options. An open question is whether their pr...
By Yu-Xiang Lin, Chen-An Li, Sheng-Lun Wei, Po-Chun Chen, Hsin-Hsi Chen, Hung-yi Lee
DiscoPhon is a multilingual benchmark designed to evaluate unsupervised phoneme discovery from discrete speech units. It includes 6 development and 6 test languages that cover a wide range of phonemic contrasts, and requires systems to generate discrete units mapped to a predefined phoneme inventory using only 10 hours of speech from an unseen language. The benchmark assesses unit quality, recognition, and segmentation, and provides four pretrained multilingual HuBERT and SpidR baselines that demonstrate current models can produce units that correlate well with phonemes, though performance varies across languages.
By Maxime Poli, Manel Khentout, Angelo Ortiz Tandazo, Ewan Dunbar, Emmanuel Chemla, Emmanuel Dupoux
arXiv:2609.15313v1 Announce Type: cross
Abstract: Autoregressive generation of interleaved text and acoustic tokens is a common approach to spoken-response generation in speech large language models....
By Daxin Tan, Dehua Tao, Chengxi Deng, Hanlin Zhang, Xiao Chen
arXiv:2607. 02633v1 Announce Type: new Abstract: We present GRAFT, a per-word pronunciation conditioning mechanism for text-to-speech neural codec language modeling.
By Antonis Asonitis, Francesco Verdini, Aref Farhadipour, Vijeta Avijeet, Pierre-Edouard Honnet, Marzieh Razavi, Juan Pablo Zuluaga Gomez