The paper introduces HealthCUES, a real‑time streaming pipeline that extracts and analyzes cough and throat‑clearing events from live spoken conversations. It detects coughs within sub‑second latency, distinguishes cough subtypes (dry, wet, barking, whooping), differentiates coughing from throat clearing, and estimates temporal boundaries, all while gating alerts based on conversational context. The system, built on Qwen3Omni, achieves high accuracy (93% F1 for cough detection) and low latency (340 ms) and has been validated by healthcare professionals for telehealth use.
By Tanmay Laud, Herprit Mahal, Subhabrata Mukherjee
arXiv:2606. 02998v1 Announce Type: new Abstract: Automated cough analysis offers a path to low-cost respiratory screening, but most existing work stops at binary COVID-19 detection.
By Nikhil Vincent
BreathGRU is a semi‑supervised Bidirectional Gated Recurrent Unit framework designed to segment speech and breath events in respiratory audio. It combines acoustic feature extraction, bidirectional recurrent modeling, pseudo‑label refinement, and duration‑constrained Segmental Viterbi decoding to produce accurate speech‑breath segmentation. In evaluations against existing methods, BreathGRU achieved the highest breath event recall, lowest onset‑localisation error, and highest Mean Match Intersection over Union, outperforming large pretrained VAD models such as Silero.
By Sania Fatima Sayed, John W. Holloway, Reyer Zwiggelaar, Faisal I. Rezwan
arXiv:2606. 10972v1 Announce Type: cross Abstract: This study aims to explore the performance of the VAR model in comparison with mel-frequency cepstral coefficient (MFCC) matrices and log-mel spectrograms using deep learning.
By Ipek Sen, Ozgur Ozdemir, Elena Battini Sonmez
The paper introduces BTS-CAFE, a federated domain generalization framework for respiratory sound classification that addresses stethoscope-induced shortcuts. It combines causality-inspired device-style interventions, counterfactual metadata augmentation, and gradient alignment to reduce style–content entanglement and promote device-invariant decision boundaries. Experiments on ICBHI and SPRSound datasets show a 3.69‑point improvement in out-of-distribution performance over the baseline and outperform conventional data augmentation and federated learning methods.
By Heejoon Koo, Yoon Tae Kim, Miika Toikkanen, June-Woo Kim
arXiv:2606. 17339v1 Announce Type: new Abstract: Speech offers a uniquely informative window into health by simultaneously engaging neurological, motor, respiratory, and vocal systems.
By Sejal Bhalla, Larry Kieu, Aina Merchant, Eyal de Lara, Alex Mariakakis
The paper introduces RALCT, a lightweight Convolutional Transformer that combines randomized audio augmentations, MFCCs, and log‑mel spectrograms to extract robust features for environmental sound recognition. With only about 310,000 parameters, RALCT achieves state‑of‑the‑art accuracy—over 93% on UrbanSound8K, peaking at 94.56%—making it suitable for deployment on mobile devices. The authors also develop a mobile app that integrates the model to provide real‑time safety alerts for hearing‑impaired users.
By Julia Huang
arXiv:2609.27222v1 Announce Type: cross
Abstract: Pneumonia is difficult to diagnose in older long-term care residents; multimorbidity and atypical presentations obscure signs, motivating operational...
By Nicholas Rasmussen, Oleg Zaslavsky, Zih-Ling Wang, Hongyu Yu, Joelle Fathi, Kaibao Nie, Amil Khanzada, Tomoko Ito
arXiv:2509. 11606v4 Announce Type: replace-cross Abstract: Cardiovascular diseases (CVDs) are the leading cause of death worldwide, accounting for approximately 17.
By Milan Marocchi, Matthew Fynn, Kayapanda Mandana, Yue Rong
This study aims to explore the performance of the VAR model in comparison with mel-frequency cepstral coefficient (MFCC) matrices and log-mel spectrograms using deep learning. In pulmonary sound classification, spectrogram-based representations suffer from inconsistent temporal dimensions due to varying respiratory cycle durations.
arXiv:2606. 27973v1 Announce Type: cross Abstract: Speech-based cognitive impairment detection offers a noninvasive, accessible alternative to costly biomarker assays, yet transformer-based models remain clinically uninterpretable.
By Yasaman Haghbin, Sina Rashidi, Ali Zolnour, Fatemeh Taherinezhad, Ali Fartoot, Hossein Azadmaleki, James M Noble, Maryam Dadkhah, Maryam Zolnoori
arXiv:2607. 10168v1 Announce Type: cross Abstract: It is still hard to find Alzheimer's disease (AD) early, especially when neuroimaging is expensive or tools that depend on language are not available.
By Rashin Gholijani Farahani, Azam Bastanfard