The study investigates why machine‑learning models for tuberculosis screening based on cough acoustics fail to generalize across datasets. Classical ML and deep‑learning classifiers achieved moderate performance within their own datasets (ROC‑AUC up to 0.755) but performed poorly on external data, often below 0.6. The authors found that audio features were more influenced by recording device and dataset than by TB status, and that device‑diverse training improved transfer while device mismatch degraded it. A clinical‑variable baseline showed more consistent generalization, suggesting acquisition‑specific variability is a stronger driver of poor generalizability than population shift.
By Wensi Zhang, Tomas Teijeiro, J\'er\^ome Thevenot, David Atienza
arXiv:2606. 02998v1 Announce Type: new Abstract: Automated cough analysis offers a path to low-cost respiratory screening, but most existing work stops at binary COVID-19 detection.
By Nikhil Vincent
arXiv:2606. 17339v1 Announce Type: new Abstract: Speech offers a uniquely informative window into health by simultaneously engaging neurological, motor, respiratory, and vocal systems.
By Sejal Bhalla, Larry Kieu, Aina Merchant, Eyal de Lara, Alex Mariakakis
arXiv:2609.16551v1 Announce Type: new
Abstract: Self-supervised learning (SSL) can reduce the need for labelled medical images, but the choice of pretext objective remains unclear for lung ultrasound...
By Moein Heidari, Junbo Rao, Jai Choraria, Wenjin Chen, David J. Foran, Ilker Hacihaliloglu
Self-supervised learning (SSL) can reduce the need for labelled medical images, but the choice of pretext objective remains unclear for lung ultrasound (LUS). Contrastive learning, masked reconstructi...
The paper introduces HealthCUES, a real‑time streaming pipeline that extracts and analyzes cough and throat‑clearing events from live spoken conversations. It detects coughs within sub‑second latency, distinguishes cough subtypes (dry, wet, barking, whooping), differentiates coughing from throat clearing, and estimates temporal boundaries, all while gating alerts based on conversational context. The system, built on Qwen3Omni, achieves high accuracy (93% F1 for cough detection) and low latency (340 ms) and has been validated by healthcare professionals for telehealth use.
By Tanmay Laud, Herprit Mahal, Subhabrata Mukherjee
The paper introduces BTS-CAFE, a federated domain generalization framework for respiratory sound classification that addresses stethoscope-induced shortcuts. It combines causality-inspired device-style interventions, counterfactual metadata augmentation, and gradient alignment to reduce style–content entanglement and promote device-invariant decision boundaries. Experiments on ICBHI and SPRSound datasets show a 3.69‑point improvement in out-of-distribution performance over the baseline and outperform conventional data augmentation and federated learning methods.
By Heejoon Koo, Yoon Tae Kim, Miika Toikkanen, June-Woo Kim
arXiv:2512.19735v4 Announce Type: replace
Abstract: Accurately predicting mortality risk in intensive care unit (ICU) patients is critical for clinical decision-making. Large language models (LLMs) a...
By Gangxiong Zhang, Yongchao Long, Yuxi Zhou, Yong Zhang, Shenda Hong
arXiv:2606. 08247v1 Announce Type: cross Abstract: Acute asthma risk assessment requires rapid interpretation of respiratory sounds, oxygenation, airflow limitation, speech ability, work of breathing, mental status, and response to reliever therapy.
By Aueaphum Aueawatthanaphisut
The study examines how three parameter‑efficient adaptation methods—linear heads on the raw CLS token, an MLP, and an attention‑pooling module—affect pathology classification accuracy and subgroup fairness when applied to a frozen Rad‑DINO chest X‑ray encoder. Using the MIMIC‑CXR dataset, the authors evaluate eight pathologies across race, sex, and imaging‑view subgroups, finding that attention pooling yields the best overall performance and encodes protected attributes most strongly, yet higher performance does not consistently reduce subgroup disparities. The results show that attribute encoding strength and layer choice do not reliably predict fairness outcomes, indicating that fairness must be assessed directly for each task.
By Dhruv Gupta, Emma A. M. Stanley, Fabio De Sousa Ribeiro, Sujal R. Desai, Ben Glocker
arXiv:2605. 00865v2 Announce Type: replace-cross Abstract: We tested whether auditory-evoked EEG supports subject-independent five-vowel perception decoding when trial identity, model identity, prediction provenance, and participant-level inference are controlled within a single benchmark.
By Xiaoyang Li, Zeyan Tao
arXiv:2607. 04526v1 Announce Type: cross Abstract: First-shot anomalous sound detection in DCASE Challenge Task 2 must flag anomalies of unseen machine types with a single threshold, without knowing whether a test clip comes from the data-rich source domain (990 normal training clips) or the data-scarce target domain (10).
By Grach Mkrtchian