arXiv:2602. 18452v3 Announce Type: replace-cross Abstract: As conversational multimodal AI tools are increasingly adopted to process patient data for health assessment, robust benchmarks are needed to measure progress and expose failure modes under realistic conditions.
By Gaia A. Bertolino, Yuwei Zhang, Tong Xia, Domenico Talia, Cecilia Mascolo
arXiv:2606. 08247v1 Announce Type: cross Abstract: Acute asthma risk assessment requires rapid interpretation of respiratory sounds, oxygenation, airflow limitation, speech ability, work of breathing, mental status, and response to reliever therapy.
By Aueaphum Aueawatthanaphisut
arXiv:2608.29241v1 Announce Type: new
Abstract: Clinical voice agents are now deployed in routine care, where real patients do not wait their turn: they interrupt. These systems typically use a casca...
By Zachary Ellis, Spencer Hazel, Adam Brandt, Yajie Vera He, Ernest Lim, Jared Joselowitz
arXiv:2608. 09861v1 Announce Type: new Abstract: Audio-visual interaction is the standard for patient-physician consultations, enabling natural communication and effective assessment of illness through non-verbal cues.
By Mahvish Nagda, Jihyeon Lee, Matthew Thompson, Chunjong Park, Tim Strother, Valentin Li\'evin, Roma Ruparel, Akshay Goel, Teya Bergamaschi, Suhana Bedi, Meet Shah, Pavel Dubov, Liviu Panait, Toshiyuki Fukuzawa, Sam Schmidgall, Craig Schiff, Joseph Xu, Aliya Rysbek, Yana Lunts, Jan Freyberg, Rebecca Hemengway, Sunny Virmani, David Racz, Carey Radebaugh, Jo\"elle Barral, Kavi Goel, Dale R. Webster, Katherine Chou, Avinatan Hassidim, Yossi Matias, James Manyika, Gregory Wayne, Tao Tu, Yun Liu, Ethan Goh, Christina Chen, Ryutaro Tanno, Po-Hsuan Cameron Chen, Mike Schaekermann, Anil Palepu
arXiv:2608.16053v2 Announce Type: replace
Abstract: Synthetic conversational speech has become an important resource for developing and evaluating conversational speech systems. However, existing dia...
By Pengcheng Wang, Sheng Li, Jiyi Li, Takahiro Shinozaki
Audio-visual interaction is the standard for patient-physician consultations, enabling natural communication and effective assessment of illness through non-verbal cues. While text-based AI has shown promise, it discards essential perceptual dimensions and limits patients who cannot articulate symptoms in writing.
VoiceLongMemEval (VLME) is a new benchmark that tests AI assistants on their ability to remember how users sounded by incorporating paralinguistic metadata—such as emotion labels, prosody descriptors, and voice events—into each conversational turn. The benchmark uses a three‑stage adversarial gate to ensure that models cannot succeed with transcript alone, revealing a significant affect gap: models gain 0.09 to 0.38 accuracy when provided with paralinguistic cues, and audio‑native models outperform standard ASR pipelines in extracting these signals. The dataset and code will be released upon acceptance.
By Ramit Pahwa, Parivesh Priye, Apoorva Beedu
arXiv:2608. 06110v1 Announce Type: new Abstract: This paper presents ECHO (Enhanced Care \& Health Observer), a locally-deployable conversational health assistant for long-term chronic care management.
By Abdulkadir K\"ul\c{c}e, Alihan Esen, Ca\u{g}la Fikir, Berke Kurt, Kuzey Arar, G\"okhan Ercan, Faik Boray Tek
arXiv:2606. 02998v1 Announce Type: new Abstract: Automated cough analysis offers a path to low-cost respiratory screening, but most existing work stops at binary COVID-19 detection.
By Nikhil Vincent
arXiv:2606. 30491v1 Announce Type: cross Abstract: Background.
By Zhuhan Bao, Rui Yang, Bohao Yang, Zhiyi Liu, Sicheng Shu, Ruio Heerschap, Le Li, Doris Yang, Elisabeth Bond, Haoyuan Wang, Nicoleta Economou-Zavlanos, Joshua M. Biro, Matthew McDermott, Nan Liu, Anand Chowdhury, Kai Sun, Kathryn Pollak, Ed Hammond, Chuan Hong
DocTalkBN is a large-scale multimodal dataset of authentic expert telemedicine conversations in Bengali, comprising 557.63 hours of paired audio and text, 1,515 multi-turn patient calls, and 10,274 host–doctor question–answer exchanges across 26 medical specialties. The dataset contains 1.7 million tokens and preserves the spontaneity and contextual richness of real medical interactions in a low-resource language. Three downstream tasks—medical triage classification, advice safety evaluation, and medical named entity recognition—are constructed to benchmark large language models and encoder-based baselines, demonstrating DocTalkBN’s practical usefulness for clinically grounded reasoning.
By Anik Saha, Fahmida Sultana Naznin, Sadatul Islam Sadi, Ananya Shahrin Promi, Wahid Al Azad Navid, Rifat Shahriyar
arXiv:2606. 05121v1 Announce Type: cross Abstract: Audio is an inherently interactive modality, yet today's Large Audio Language Models (LALMs) are offline, and streaming audio models each handle only a single task such as streaming ASR or voice chatting.
By Zhifei Xie, Zihang Liu, Ze An, Xiaobin Hu, Yue Liao, Ziyang Ma, Dongchao Yang, Mingbao Lin, Deheng Ye, Shuicheng Yan, Chunyan Miao