SonicCaps is a large-scale audio captioning dataset featuring approximately 15 million captions paired with 700,000 audio clips, created using the Qwen3-Omni multimodal language model. The dataset emphasizes diversity by generating around 24 captions per clip through structured prompt engineering and few-shot generation, covering main descriptions, rephrased variants, and semantic tags. Human evaluations rate SonicCaps higher than existing datasets, and training CLAP models on it improves audio retrieval and zero-shot classification across public and commercial benchmarks.
By Zineb Lahrichi, Marc Ferras, Ga\"el Richard, Geoffroy Peeters
arXiv:2607. 27109v2 Announce Type: cross Abstract: With the development of audio large language models (AudioLLMs), audio captioning needs to move from brief descriptions toward open-ended and fine-grained free-form descriptions.
By Weijie Wu, Junbo Li, Lin Li, Jun Fang, Qingyang Hong
The paper introduces STAG, a post‑hoc framework that provides token‑level spectro‑temporal grounding for captions produced by audio‑based multimodal large language models (MLLMs). STAG estimates temporal support for each token via vocabulary projections of encoded audio, measures frequency‑band relevance through controlled spectral occlusion, and fuses these signals into a spectro‑temporal relevance map. Evaluations across ten explanation methods and four grounding benchmarks show that STAG achieves superior event‑localization performance on every dataset, and counterfactual deletion experiments confirm that removing the identified evidence selectively reduces model confidence and often eliminates the corresponding event from regenerated captions.
By Lucia Cascone, Valeria Fraenza, Michele Nappi, Fabio Narducci, Benedetto Simone
arXiv:2608. 04479v1 Announce Type: cross Abstract: Text-to-audio (TTA) generation has recently achieved remarkable progress in synthesizing realistic audio from natural language descriptions.
By Jinting Wang, Yuguang Yang, Shengyu Li, Yan Rong, Shan Yang, Xiaoda Yang, Li Liu
arXiv:2606. 07387v1 Announce Type: new Abstract: State-of-the-art text-to-music generation systems rely on massive proprietary datasets and industrial-scale compute, making it impossible to disentangle architectural contributions from resource advantages.
By Yun-Chen Cheng, Tzu-Hung Huang, Chih-Pin Tan
arXiv:2511. 16757v2 Announce Type: replace-cross Abstract: Audio-language pretraining (ALP) holds promise for learning general-purpose audio representation, yet remains underexplored.
By Wei-Cheng Tseng, Xuanru Zhou, Mingyue Huo, Yiwen Shao, Hao Zhang, Dong Yu