The paper introduces HealthCUES, a real‑time streaming pipeline that extracts and analyzes cough and throat‑clearing events from live spoken conversations. It detects coughs within sub‑second latency, distinguishes cough subtypes (dry, wet, barking, whooping), differentiates coughing from throat clearing, and estimates temporal boundaries, all while gating alerts based on conversational context. The system, built on Qwen3Omni, achieves high accuracy (93% F1 for cough detection) and low latency (340 ms) and has been validated by healthcare professionals for telehealth use.
By Tanmay Laud, Herprit Mahal, Subhabrata Mukherjee
arXiv:2606. 02998v1 Announce Type: new Abstract: Automated cough analysis offers a path to low-cost respiratory screening, but most existing work stops at binary COVID-19 detection.
By Nikhil Vincent
BreathGRU is a semi‑supervised Bidirectional Gated Recurrent Unit framework designed to segment speech and breath events in respiratory audio. It combines acoustic feature extraction, bidirectional recurrent modeling, pseudo‑label refinement, and duration‑constrained Segmental Viterbi decoding to produce accurate speech‑breath segmentation. In evaluations against existing methods, BreathGRU achieved the highest breath event recall, lowest onset‑localisation error, and highest Mean Match Intersection over Union, outperforming large pretrained VAD models such as Silero.
By Sania Fatima Sayed, John W. Holloway, Reyer Zwiggelaar, Faisal I. Rezwan
arXiv:2606. 10972v1 Announce Type: cross Abstract: This study aims to explore the performance of the VAR model in comparison with mel-frequency cepstral coefficient (MFCC) matrices and log-mel spectrograms using deep learning.
By Ipek Sen, Ozgur Ozdemir, Elena Battini Sonmez
The paper introduces BTS-CAFE, a federated domain generalization framework for respiratory sound classification that addresses stethoscope-induced shortcuts. It combines causality-inspired device-style interventions, counterfactual metadata augmentation, and gradient alignment to reduce style–content entanglement and promote device-invariant decision boundaries. Experiments on ICBHI and SPRSound datasets show a 3.69‑point improvement in out-of-distribution performance over the baseline and outperform conventional data augmentation and federated learning methods.
By Heejoon Koo, Yoon Tae Kim, Miika Toikkanen, June-Woo Kim
arXiv:2606. 17339v1 Announce Type: new Abstract: Speech offers a uniquely informative window into health by simultaneously engaging neurological, motor, respiratory, and vocal systems.
By Sejal Bhalla, Larry Kieu, Aina Merchant, Eyal de Lara, Alex Mariakakis