arXiv:2609.01246v1 Announce Type: new
Abstract: Current Large Language Models (LLMs) are primarily optimized for written text, often producing outputs that are grammatically correct and helpful yet p...
By Thibaut Thonet, Jos Rozen, Laurent Besacier
arXiv:2609.35952v1 Announce Type: cross
Abstract: We introduce HEAR (Human-recorded Evaluation of Audio-LLM bias by Real speakers), a large-scale, ecologically valid benchmark comprising 87k real hum...
By Shen Yan, Duc Le, Irina-Elena Veliche
arXiv:2603. 09714v2 Announce Type: replace-cross Abstract: While multi-audio understanding is critical for large audio-language models (LALMs), it remains underexplored.
By Chih-Kai Yang, Yun-Shao Tsai, Yu-Kai Guo, Ping-Le Tsai, Yen-Ting Piao, Hung-Wei Chen, Ting-Lin Hsiao, Yun-Man Hsu, Ke-Han Lu, Hung-yi Lee
Large Audio-Language Models (LALMs) have been widely used as judge models for the automatic evaluation of generated speech. However, prior approaches predominantly focus on holistic naturalness, leaving fine-grained paralinguistic distinctions underexplored.
arXiv:2608. 13624v1 Announce Type: cross Abstract: Large Audio Language Models (LALMs) have seen increasing use for audio understanding tasks such as speech recognition and audio question answering, raising concerns about fairness across demographic subgroups.
By Zhe Liu
arXiv:1807. 08636v2 Announce Type: replace-cross Abstract: In music and audio production, attenuation of spectral resonances is an important step towards a technically correct result.
By Maarten Grachten, Emmanuel Deruty, Alexandre Tanguy
Modern TTS systems approach human quality for high-resource languages but degrade when clean speech data is scarce. Modern Greek exemplifies this, lacking the curated corpora behind state-of-the-art s...
arXiv:2609.10022v1 Announce Type: cross
Abstract: Modern TTS systems approach human quality for high-resource languages but degrade when clean speech data is scarce. Modern Greek exemplifies this, la...
By Georgios Syllas, Efthymios Georgiou, Kosmas Kritsis, Alexandros Potamianos
arXiv:2605.00022v2 Announce Type: replace-cross
Abstract: The rapid proliferation of large audio models (LAMs) demands efficient approaches for model comparison, yet comprehensive benchmarks are cost...
By Woody Haosheng Gan, William Held, Diyi Yang
The paper investigates how Spoken Language Models (SLMs) process speech compared to text, noting that current SLMs show weak alignment between speech and text representations despite strong downstream performance. The authors propose a framework that separates length mismatch from semantic alignment to better match speech and text representations. Experiments on multiple benchmarks demonstrate that this approach yields competitive results against strong baselines, highlighting the need to explicitly address structural differences between speech and text in SLM training.
By Hyeonyu Kim, Hwayeon Kim, Youngwon Choi, Myeongkyun Cho, Huu-Kim Nguyen
AudioLLMs enable speech recognition conditioned on textual prompts such as domain descriptions or entity lists. However, it remains unclear whether these models genuinely utilise such context or rely on parametric knowledge learned during pretraining.
The paper evaluates how robust three text‑to‑audio models—MusicGen‑small, MusicGen‑large, and Stable Audio 2.5—are to small changes in prompts that could affect adaptive game soundtracks. Using metrics such as log‑Mel distance, MFCC/chroma‑DTW, and CLAP similarity, the study finds that Stable Audio 2.5 consistently yields the lowest acoustic distances and highest CLAP similarity when prompts are structurally rephrased, while MusicGen‑large performs best under lexical substitutions and intensity shifts. The authors also observe that Stable Audio 2.5 shows the greatest variation in prompt‑to‑audio alignment across different random seeds, highlighting the need for multi‑seed robustness testing in game audio applications.
By Jiahui Wu, Mei Si