The paper investigates whether audio language models encode phonetic features similarly when processing spoken versus written input. By comparing mean representations of minimal phoneme pairs across six models, seven features, and 15 languages, the study finds that only voicing in two Qwen2.5-Omni models shows a significant shared direction, and that the model family—not size—determines feature representation. The analysis uses cosine similarity against a random-pair reference to assess alignment across modalities.
The paper investigates whether audio‑language models capture paralinguistic cues beyond spoken content. Using the Expresso dataset and four open‑source models, the authors trace how speaking style information is encoded in the late layers of the audio encoder but is degraded before reaching the final output. They find that some models are content‑driven while others are acoustic‑driven, revealing a gap between what is encoded and what is utilized in current audio‑language models.
By Bhuvan Koduru, Dareen Safar B Alharthi, Rita Singh, Bhiksha Raj
arXiv:2603. 09714v2 Announce Type: replace-cross Abstract: While multi-audio understanding is critical for large audio-language models (LALMs), it remains underexplored.
By Chih-Kai Yang, Yun-Shao Tsai, Yu-Kai Guo, Ping-Le Tsai, Yen-Ting Piao, Hung-Wei Chen, Ting-Lin Hsiao, Yun-Man Hsu, Ke-Han Lu, Hung-yi Lee
The paper introduces an Encoding Probe that reconstructs language model representations using interpretable features, addressing limitations of traditional decoding probes such as incomparable feature contributions and correlation effects. It evaluates this approach on text and speech transformer models, examining features from acoustics, phonetics, syntax, lexicon, and speaker identity. Findings reveal that speaker-related effects vary with training objectives and datasets, while syntactic and lexical features independently contribute to reconstruction, offering a complementary perspective on model interpretation.
By Gaofei Shen, Martijn Bentum, Tomas O. Lentz, Afra Alishahi, Grzegorz Chrupa{\l}a
arXiv:2603. 28378v2 Announce Type: replace-cross Abstract: We present the first systematic Membership Inference Attack (MIA) evaluation of LALMs.
By Jia-Kai Dong, Yu-Xiang Lin, Hung-Yi Lee
arXiv:2609.30483v1 Announce Type: cross
Abstract: Audio language models state numbers for acoustic quantities, and neither human opinion nor a judge model says whether such a number is true of the si...
By Sheng-Tse Lin, Siyuan Zhai, Chien-Liang Kuo, Massa Baali, Bhiksha Raj