The paper introduces HealthCUES, a real‑time streaming pipeline that extracts and analyzes cough and throat‑clearing events from live spoken conversations. It detects coughs within sub‑second latency, distinguishes cough subtypes (dry, wet, barking, whooping), differentiates coughing from throat clearing, and estimates temporal boundaries, all while gating alerts based on conversational context. The system, built on Qwen3Omni, achieves high accuracy (93% F1 for cough detection) and low latency (340 ms) and has been validated by healthcare professionals for telehealth use.
By Tanmay Laud, Herprit Mahal, Subhabrata Mukherjee
arXiv:2606. 17339v1 Announce Type: new Abstract: Speech offers a uniquely informative window into health by simultaneously engaging neurological, motor, respiratory, and vocal systems.
By Sejal Bhalla, Larry Kieu, Aina Merchant, Eyal de Lara, Alex Mariakakis
arXiv:2609.22452v1 Announce Type: new
Abstract: Long-context understanding remains a fundamental challenge for large language models, as excessively long inputs often lead models to forget salient in...
By Xize Cheng, Wenxu Jia, Chenyuhao Wen, Dongjie Fu, Zehan Wang, Xinyu Zhang, Tao Jin
arXiv:2609.22214v1 Announce Type: new
Abstract: Long multilingual conversational spoken question answering requires systems to balance long-range transcript semantics with sparse acoustic and speaker...
By Shangkun Huang, Junchao Hu, Huan Shen, Guoji Wang, Yingao Wang, Shaosai Li, Wei Zou, Yunzhang Chen
SONIC‑O1 is a new benchmark designed to evaluate multimodal large language models on audio‑video understanding. It contains 60 hours of 231 clips across 13 real‑world conversational domains, with 4,958 human‑verified annotations and demographic metadata. The benchmark tests open‑ended summarization, multiple‑choice question answering, and temporally grounded reasoning, revealing performance gaps between model families and across demographic groups.
By Ahmed Y. Radwan, Christos Emmanouilidis, Hina Tabassum, Deval Pandya, Shaina Raza
EXAM$^2$ is a new benchmark for audio understanding that covers six languages and multiple modalities—speech, sound, music, mixed-audio, and visual images—providing 5,667 multiple-choice questions, 22,614 image instances, and 135,684 multilingual translations. It evaluates large audio language models (LALMs) and multimodal large language models (LLMs), revealing significant gaps in multilingual and cross‑modal performance. The authors also introduce Gemma3n-EXAM$^2$, a lightweight fusion model that improves multilingual results by up to 12.4% and multimodal results by 21.7% over a strong baseline.
By Jiawen Wang, Xiaoxue Gao, Zi Haur Pang, Nancy F. Chen
arXiv:2609.35952v1 Announce Type: cross
Abstract: We introduce HEAR (Human-recorded Evaluation of Audio-LLM bias by Real speakers), a large-scale, ecologically valid benchmark comprising 87k real hum...
By Shen Yan, Duc Le, Irina-Elena Veliche
arXiv:2602. 14612v4 Announce Type: replace-cross Abstract: Answering natural-language questions over multi-hour audio requires both event recognition and temporal grounding.
By Kartik Hegde, Arvind Krishna Sridhar, Naveen Vakada, Yinyi Guo, Erik Visser
DocTalkBN is a large-scale multimodal dataset of authentic expert telemedicine conversations in Bengali, comprising 557.63 hours of paired audio and text, 1,515 multi-turn patient calls, and 10,274 host–doctor question–answer exchanges across 26 medical specialties. The dataset contains 1.7 million tokens and preserves the spontaneity and contextual richness of real medical interactions in a low-resource language. Three downstream tasks—medical triage classification, advice safety evaluation, and medical named entity recognition—are constructed to benchmark large language models and encoder-based baselines, demonstrating DocTalkBN’s practical usefulness for clinically grounded reasoning.
By Anik Saha, Fahmida Sultana Naznin, Sadatul Islam Sadi, Ananya Shahrin Promi, Wahid Al Azad Navid, Rifat Shahriyar
arXiv:2609.23416v1 Announce Type: cross
Abstract: Long-form audio performance is often summarized by context length and aggregate accuracy, obscuring how language, evidence, and task jointly shape di...
By Zeyu Yang, Xinyu Zhang, Zibo Bi, Pei Zhang, Xize Cheng, Jin Xu, Baosong Yang, Satoshi Nakamura
MMTClinic is a new benchmark that tests large language models on complex reasoning and question‑answering tasks involving clinical time‑series data. It combines text, medical images, and multivariate physiological signals to create 30,000 QA pairs—including 15,000 multiple‑choice and 15,000 open‑ended questions—in five languages (English, Hindi, Bengali, Marathi, and Tamil). The benchmark covers mortality prediction, heart‑rate forecasting, and SOFA score estimation, and evaluates 13 state‑of‑the‑art LLMs across zero‑shot, few‑shot, and chain‑of‑thought settings, revealing significant performance gaps across tasks, languages, and modalities.
By Sourav Malakar, Harshit Nigam, Akash Ghosh, Sriparna Saha, Amlan Chakrabarti, Saptarsi Goswami, Priti Singh
arXiv:2609.22771v1 Announce Type: cross
Abstract: Recent advances in Audio LLMs have achieved human-level speech recognition, yet existing systems struggle to capture paralinguistic aspects such as s...
By Nishit Anand, Jiaqi Su, Ke Chen, Yunyun Wang, Dinesh Manocha, Ramani Duraiswami, Rithesh Kumar, Zeyu Jin