arXiv:2607. 04383v1 Announce Type: cross Abstract: Large Audio-Language Models (LALMs) reason fluently about sound yet struggle to localize precisely when events occur, while classical Sound Event Detection attains frame-level precision only over a closed label set.
By Zihan Zhang, Xize Cheng, Wenhao Yan, Tong Zhang, Dongjie Fu, Boyun Zhang, Yongbo He, Tao Jin
arXiv:2609.15215v1 Announce Type: cross
Abstract: Large Audio-Language Models (LALMs) have substantially advanced general audio understanding, yet they remain limited in fine-grained temporal percept...
By Yanfeng Shi, Yan Song, Junhui Li, Tinggan Huang, Wu Guo, Haoyu Song, Ian McLoughlin
arXiv:2606. 01802v1 Announce Type: cross Abstract: MOSS-Audio is a unified audio-language model for speech, environmental sound, and music understanding, supporting audio captioning, time-aware question answering, timestamped transcription, and audio-grounded reasoning.
By Chen Yang, Chufan Yu, Hanfu Chen, Jie Zhu, Jingqi Chen, Ke Chen, Wenxuan Wang, Yang Wang, Yaozhou Jiang, Yi Jiang, Zhengyuan Lin, Ziqi Chen, Zhaoye Fei, Chenghao Liu, Jun Zhan, Kang Yu, Kexin Huang, Mingshu Chen, Qinyuan Cheng, Ruixiao Li, Shimin Li, Songlin Wang, Yang Gao, Yiyang Zhang, Xipeng Qiu
TEMPO is a unified model that adds temporally‑grounded capabilities to large audio‑language models, enabling timestamping of events, speakers, and sounds in audio, speech, and music. It introduces a supervised fine‑tuning stage featuring atomic timestamp tokens, a time‑aware projector with sinusoidal encodings, and a distance‑aware Gaussian loss, trained via a synthetic‑to‑real curriculum. Additionally, TEMPO employs reinforcement learning (GRPO) as a refinement step, and achieves state‑of‑the‑art performance on a benchmark of 10K samples across five timestamping tasks, surpassing Audio Flamingo Next and Qwen3‑Omni.
By Apoorva Kulkarni, Kaousheik Jayakumar, Sreyan Ghosh, Utathya Aich, Ramani Duraiswami, Dinesh Manocha
ReasonAudio is a new benchmark designed to evaluate reasoning capabilities in text‑audio retrieval, addressing the gap left by existing semantic‑matching focused datasets. It tests four logical abilities—negation, temporal order, sound co‑occurrence, and sound duration—across five synthetic subtasks (1,000 queries over 10,000 composite clips) and one natural subtask (100 queries over 1,000 real‑world clips). Evaluation of 11 state‑of‑the‑art systems shows significant limitations, with the best model, OmniEmbed‑7B, scoring only 20.7 overall and 53.8% in a controlled setting, compared to 70.6% for its generative backbone and 95.6% for humans.
By Honglei Zhang, Yuting Chen, Chenpeng Hu, Pengfei Zhou, Siyue Zhang, Yilei Shi
arXiv:2609.22452v1 Announce Type: new
Abstract: Long-context understanding remains a fundamental challenge for large language models, as excessively long inputs often lead models to forget salient in...
By Xize Cheng, Wenxu Jia, Chenyuhao Wen, Dongjie Fu, Zehan Wang, Xinyu Zhang, Tao Jin
MusTBench is a music‑expert‑validated benchmark that evaluates temporal grounding in Large Audio‑Language Models (LALMs) through five temporally grounded question‑answering tasks. The paper also introduces MusT, a four‑stage optimization recipe—music encoder adaptation, LLM adaptation, supervised fine‑tuning, and RL‑based optimization—to improve temporal grounding. Experiments show that current LALMs struggle with precise temporal grounding, while MusT yields significant improvements, highlighting temporal grounding as a key missing capability in these models.
By Daeyong Kwon, Qiyu Wu, Shinobu Kuriya, Junghyun Koo, Shuyang Cui, Zhi Zhong, Wei-Hsiang Liao, Hiromi Wakaki, Yuki Mitsufuji
LEAP is a framework for long audio‑video question answering that avoids encoding entire recordings by dividing them into fixed‑duration blocks. It performs a lightweight localization pass on each block to score short candidate windows, then pools the highest‑ranked windows for a single bounded answer pass, keeping the answer input and peak context independent of recording length. The method trains both a localization LoRA and an answer LoRA, supports causal streaming queries, and achieves significant performance gains over baseline models on multiple AVQA benchmarks.
By Juyi Lin, Zhiqiang Lao, Jiali Cui, Lin Zhao, Pu Zhao, Dichang Zhang, Arman Akbari, Yu Qi, Xinru Jiang, Yanzhi Wang, Heather Yu, Liang Peng
arXiv:2606. 12300v1 Announce Type: cross Abstract: Temporal grounding--returning the interval $[t_s, t_e]$ for a natural-language query over a video--is the language interface to long-form video, yet has been studied on short videos; the dynamics of hour-scale natural-language grounding remain underexplored.
By Sukmin Seo, Geewook Kim
arXiv:2607. 16736v1 Announce Type: cross Abstract: This paper presents RealDESED, a real-world domestic sound event detection (SED) benchmark comprising 5,710 audio recordings collected by 652 participants in their homes.
By Florian Schmid, Paul Primus, Alexander Fichtinger, Tara Jadidi, Tobias Morocutti, Gerhard Widmer
arXiv:2603. 18558v2 Announce Type: replace-cross Abstract: Long-form video question answering requires reasoning over extended temporal contexts, making frame selection a critical bottleneck for multi-modal large language models (MLLMs) bound by finite context windows.
By Dan Ben-Ami, Gabriele Serussi, Kobi Cohen, Chaim Baskin
VoiceTrace introduces a new benchmark, VoiceTrace-Bench, for hybrid speech retrieval that combines a textual query specifying "what" to retrieve with a reference speech specifying "who" to retrieve. The authors propose a two‑stage framework: VoiceTrace‑Emb, which learns unified audio‑text embeddings for efficient large‑scale retrieval, and VoiceTrace‑Reranker, which fine‑grains relevance by jointly examining query‑candidate pairs. Experiments show VoiceTrace outperforms existing methods on both traditional semantic speech retrieval benchmarks and the new hybrid setting.
By Aaron Yee, Fengjie Lu, Jiarui Hai, Chenang Jiang, Helin Wang, Siwei Tu, Weitao You, Lingyun Sun