arXiv:2608. 19211v1 Announce Type: cross Abstract: Human speech is richly expressive, with prosody carrying linguistic and emotional information beyond the lexical content.
By Linkai Peng, Baorian Nuchged
arXiv:2609.05871v1 Announce Type: cross
Abstract: Audio-conditioned language models often underuse acoustic cues such as prosody, emotion, and non-speech sounds, raising the question of whether ASR-s...
By Song-ha Jo, Sehyun Lee, Soyoon Kim, Jaesik Choi, Sanghyuk Choi
The study investigates how safety alignment in large language models, trained mainly in English, transfers to other languages. While models show near-perfect harmfulness detection (AUROC > 0.98) using unrelated harmless prompts (easy negatives), performance drops sharply in low‑resource languages when using surface‑similar benign prompts (hard negatives). This degradation persists across multiple languages and models, indicating that easy‑negative evaluation alone cannot confirm cross‑lingual harmfulness representation quality.
By Paras Balani, Subhrakanta Panda
arXiv:2607. 14147v1 Announce Type: cross Abstract: Aligned language models refuse harmful requests, but a one-line prefill ("Sure, here is") strips the refusal.
By Alex Kwon
The study investigates the internal workings of an audio language model (Qwen3-Omni) by applying a logit lens to its middle layers. It finds that the model’s reasoning about spoken questions becomes legible in words before any token is emitted, revealing language‑agnostic, paralinguistic, and temporally distinct signals that are causally used in the network’s decision process. The authors demonstrate that these signals can be isolated and mapped to specific layers, providing a qualitative account of how the model processes audio input.
By Jiajun Fan, Jingyuan Li, Prashanth Gurunath Shivakumar, Qi Luo, Jia-Hong Huang, M. Maruf, Roger Ren, Yile Gu, Rahul Pandey, Ge Liu, Ivan Bulyko
The paper presents the first music‑specific, layer‑wise empirical study of hallucination in audio‑language models, framing it as a hierarchical perceptual grounding failure across five layers: sound events, temporal properties, tonal attributes, style, and emotion. It introduces MuseDiag, a diagnostic framework that evaluates nine models and finds universal vocal misperception, significant tonal perception differences, and identifies Audio‑Flamingo‑3 as the most stable model. The study also proposes two training‑free mitigation methods, ADD‑M and TPA, which reduce hallucination in probing but show variable effectiveness in free‑form generation, highlighting the need for multi‑paradigm evaluation.
By Yu Liu, Jiahui Liu, Zhilin Liu, Cong Cao, Fangfang Yuan, Yuling Yang, Pin Xu, Yanbing Liu