arXiv:2511. 16757v2 Announce Type: replace-cross Abstract: Audio-language pretraining (ALP) holds promise for learning general-purpose audio representation, yet remains underexplored.
By Wei-Cheng Tseng, Xuanru Zhou, Mingyue Huo, Yiwen Shao, Hao Zhang, Dong Yu
arXiv:2607. 27109v2 Announce Type: cross Abstract: With the development of audio large language models (AudioLLMs), audio captioning needs to move from brief descriptions toward open-ended and fine-grained free-form descriptions.
By Weijie Wu, Junbo Li, Lin Li, Jun Fang, Qingyang Hong
arXiv:2606. 06615v1 Announce Type: cross Abstract: Retrieving music using natural language descriptions has improved with contrastive audio-text models such as CLAP, but current systems remain limited to coarse semantic queries.
By Nishit Anand, Ashish Seth, Sreyan Ghosh, Dinesh Manocha, Ramani Duraiswami
arXiv:2607. 10299v1 Announce Type: new Abstract: Recent advances in large-scale multimodal models have drivenremarkable progress in vision-language tasks; however, comprehensiveomni-modal understanding remains under-explored, largely due to thescarcity of datasets with rich, explicitly aligned auditory cues.
By Kaiying Yan, Luoyi Sun, Xiao Zhou, Weidi Xie
arXiv:2606. 14591v1 Announce Type: cross Abstract: Large Audio-Language Models (LALMs) have shown strong performance on a wide range of audio understanding tasks, yet they still struggle with complex audio reasoning.
By Hui Geng, Yi Su, Han Yin, Tianjiao Wan, Qisheng Xu, Jiaxin Chen, Zijian Gao, Hengzhu Liu, Xie Chen, Kele Xu
arXiv:2502. 16584v2 Announce Type: replace-cross Abstract: Recent advancements in audio tokenization have significantly enhanced the integration of audio capabilities into large language models (LLMs).
By Liumeng Xue, Ziya Zhou, Jiahao Pan, Zixuan Li, Shuai Fan, Yinghao Ma, Sitong Cheng, Dongchao Yang, Haohan Guo, Yujia Xiao, Xinsheng Wang, Zixuan Shen, Chuanbo Zhu, Xinshen Zhang, Tianchi Liu, Ruibin Yuan, Zeyue Tian, Haohe Liu, Xingjian Du, Emmanouil Benetos, Ge Zhang, Yike Guo, Wei Xue
The paper introduces STAG, a post‑hoc framework that provides token‑level spectro‑temporal grounding for captions produced by audio‑based multimodal large language models (MLLMs). STAG estimates temporal support for each token via vocabulary projections of encoded audio, measures frequency‑band relevance through controlled spectral occlusion, and fuses these signals into a spectro‑temporal relevance map. Evaluations across ten explanation methods and four grounding benchmarks show that STAG achieves superior event‑localization performance on every dataset, and counterfactual deletion experiments confirm that removing the identified evidence selectively reduces model confidence and often eliminates the corresponding event from regenerated captions.
By Lucia Cascone, Valeria Fraenza, Michele Nappi, Fabio Narducci, Benedetto Simone
arXiv:2609.09973v1 Announce Type: new
Abstract: Evaluating video captioning remains a critical challenge for Visual Large Language Models (VLLMs). Existing metrics primarily rely on matching generate...
By Zizhen Wang, Bo Feng, Zhengfeng Lai, Shiyu Li, Yang Lu, Meng Cao, Ping Huang, Xiaoming Simon Wang
arXiv:2509. 14659v3 Announce Type: replace-cross Abstract: Current audio captioning relies on supervised learning with paired audio-caption data, which is costly to curate and may not reflect human preferences in real-world scenarios.
By Kartik Hegde, Rehana Mahfuz, Yinyi Guo, Erik Visser
arXiv:2606. 01802v1 Announce Type: cross Abstract: MOSS-Audio is a unified audio-language model for speech, environmental sound, and music understanding, supporting audio captioning, time-aware question answering, timestamped transcription, and audio-grounded reasoning.
By Chen Yang, Chufan Yu, Hanfu Chen, Jie Zhu, Jingqi Chen, Ke Chen, Wenxuan Wang, Yang Wang, Yaozhou Jiang, Yi Jiang, Zhengyuan Lin, Ziqi Chen, Zhaoye Fei, Chenghao Liu, Jun Zhan, Kang Yu, Kexin Huang, Mingshu Chen, Qinyuan Cheng, Ruixiao Li, Shimin Li, Songlin Wang, Yang Gao, Yiyang Zhang, Xipeng Qiu
arXiv:2606. 30682v1 Announce Type: cross Abstract: Recent advances in language--audio retrieval have been largely driven by contrastive dual-encoder architectures that align audio and text in a shared embedding space.
By Fengjie Lu, Chenang Jiang, Jiarui Hai, Helin Wang, Aaron Yee
The paper investigates whether joint language‑audio embedding models encode human perceptual timbre semantics. It evaluates several state‑of‑the‑art models, finding that LAION‑CLAP aligns best with human‑perceived timbre across instrumental sounds and descriptor‑conditioned audio effects, yet the overall alignment remains limited. The study also notes that reverb‑induced timbre semantics are more consistently captured than equalization‑induced ones.
By Qixin Deng, Bryan Pardo, Thrasyvoulos N Pappas