arXiv Machine Learning

CoughSense: Five-Class Respiratory Disease Classification via Whisper Encoder Fine-Tuning and Dual-Encoder Cross-Attention Fusion with Balanced Contrastive Learning

arXiv:2606. 02998v1 Announce Type: new Abstract: Automated cough analysis offers a path to low-cost respiratory screening, but most existing work stops at binary COVID-19 detection.

arXiv Machine Learning
Jul 27

Autoregressive EHR Foundation Models with Multimodal Inputs

arXiv:2607. 22264v1 Announce Type: new Abstract: Autoregressive foundation models trained on tokenized electronic health records (EHRs) can support zero-shot clinical prediction, yet most operate on structured event codes alone, and do not incorporate multiple modalities in a principled way.

By Yuxuan Liu, Joshua Placidi, Jinpei Han, Alfred John Balston, Marek Rei, A. Aldo Faisal
arXiv AI
Jun 9

AeroSpectra Sentinel: An Auditable LLM Prompt-Chaining Decision-Support Workflow for Acute Asthma Risk Assessment from Respiratory Sounds and Clinical Signals

arXiv:2606. 08247v1 Announce Type: cross Abstract: Acute asthma risk assessment requires rapid interpretation of respiratory sounds, oxygenation, airflow limitation, speech ability, work of breathing, mental status, and response to reliever therapy.

By Aueaphum Aueawatthanaphisut
arXiv AI
Aug 28

From Sound to Symptom: Real-Time Respiratory Signal Understanding for Conversational Healthcare Agents

The paper introduces HealthCUES, a real‑time streaming pipeline that extracts and analyzes cough and throat‑clearing events from live spoken conversations. It detects coughs within sub‑second latency, distinguishes cough subtypes (dry, wet, barking, whooping), differentiates coughing from throat clearing, and estimates temporal boundaries, all while gating alerts based on conversational context. The system, built on Qwen3Omni, achieves high accuracy (93% F1 for cough detection) and low latency (340 ms) and has been validated by healthcare professionals for telehealth use.

By Tanmay Laud, Herprit Mahal, Subhabrata Mukherjee
arXiv Computer Vision
Aug 25

Extending the Horizon of Early Diagnosis: Lung Cancer Prediction with Vision Transformers

arXiv:2608.21571v1 Announce Type: new Abstract: Lung cancer remains a leading cause of cancer-related mortality worldwide, and early diagnosis is critical for improving survival. However, early-stage...

By Olivera Kotevska, Ian Goethert, Michael McGee, Maria Mahbub, Sean R. Wilkinson, Rowena Yip, Myvizhi Esai Selvan, Zeynep H. Gumus, Claudia Henschke, Robert J. Klein, Providencia Morales, Samuel M Aguayo, Ioana Danciu, Mayanka Chandrashekar
arXiv AI
Sep 2

Zero-Shot Respiratory Sound Classification through LLM-Augmented Audio-Text Alignment

The paper introduces a framework that aligns self‑supervised respiratory encoders with medical terminology in a shared latent space, enabling zero‑shot respiratory sound classification. By using a medical LLM to generate structured reports from metadata, the method creates dense semantic anchors for contrastive learning, combining a sigmoid‑based contrastive loss with the encoder’s native SSL objective and similarity‑aware negative sampling. On nine tasks across six datasets, the approach achieves a 61.3% mean zero‑shot AUC, outperforming CLAP and Qwen2‑Audio, and reaches the highest linear probing AUC with only 43% of the data used by full‑scale baselines.

By Mustafa Talha \.Ilerisoy, Hung Manh Pham, Mathias Funk, Mykola Pechenizkiy, Aaqib Saeed
arXiv AI
Jun 16

Beyond Classification: A Cough Regression Benchmark for Respiratory Acoustic Foundation Models

arXiv:2606. 15436v1 Announce Type: cross Abstract: Respiratory acoustic foundation models (FMs) excel at cough classification, yet their ability to predict continuous health quantities from cough audio remains largely unexplored, despite the clinical value of passive age, BMI, and disease probability estimation in settings where physical measurements are unavailable.

By Mayur Sanap, Prasanna Desikan, Edgar Lobaton
arXiv AI
Sep 7

Cross-modal triage network: a multimodal deep learning framework for severity-based triage and visual explainability in chest radiographs

The paper introduces the Cross‑Modal Triage Network (CMTN), a multimodal deep‑learning model that fuses a Swin Transformer V2 visual encoder with a PubMedBERT text encoder to perform severity‑based triage, pathology detection, and generate visual explanations for chest radiographs. Trained on 34,639 image‑text pairs from MIMIC‑CXR‑JPG, the CMTN achieves high ordinal agreement with reference labels (QWK = 0.9341) and excellent pathology detection (macro‑AUROC = 0.9970) while operating with 34 ms latency. However, a blinded clinical audit revealed low agreement with expert radiologists (QWK = 0.1399) and only modest spatial‑semantic concordance in heatmaps, underscoring the gap between algorithmic performance and clinical judgment.

By Zinah Ghulam, Richa Mittal, Eranga Ukwatta