arXiv:2607. 22264v1 Announce Type: new Abstract: Autoregressive foundation models trained on tokenized electronic health records (EHRs) can support zero-shot clinical prediction, yet most operate on structured event codes alone, and do not incorporate multiple modalities in a principled way.
By Yuxuan Liu, Joshua Placidi, Jinpei Han, Alfred John Balston, Marek Rei, A. Aldo Faisal
arXiv:2606. 08247v1 Announce Type: cross Abstract: Acute asthma risk assessment requires rapid interpretation of respiratory sounds, oxygenation, airflow limitation, speech ability, work of breathing, mental status, and response to reliever therapy.
By Aueaphum Aueawatthanaphisut
The paper introduces HealthCUES, a real‑time streaming pipeline that extracts and analyzes cough and throat‑clearing events from live spoken conversations. It detects coughs within sub‑second latency, distinguishes cough subtypes (dry, wet, barking, whooping), differentiates coughing from throat clearing, and estimates temporal boundaries, all while gating alerts based on conversational context. The system, built on Qwen3Omni, achieves high accuracy (93% F1 for cough detection) and low latency (340 ms) and has been validated by healthcare professionals for telehealth use.
By Tanmay Laud, Herprit Mahal, Subhabrata Mukherjee
arXiv:2608.21571v1 Announce Type: new
Abstract: Lung cancer remains a leading cause of cancer-related mortality worldwide, and early diagnosis is critical for improving survival. However, early-stage...
By Olivera Kotevska, Ian Goethert, Michael McGee, Maria Mahbub, Sean R. Wilkinson, Rowena Yip, Myvizhi Esai Selvan, Zeynep H. Gumus, Claudia Henschke, Robert J. Klein, Providencia Morales, Samuel M Aguayo, Ioana Danciu, Mayanka Chandrashekar
The paper introduces a framework that aligns self‑supervised respiratory encoders with medical terminology in a shared latent space, enabling zero‑shot respiratory sound classification. By using a medical LLM to generate structured reports from metadata, the method creates dense semantic anchors for contrastive learning, combining a sigmoid‑based contrastive loss with the encoder’s native SSL objective and similarity‑aware negative sampling. On nine tasks across six datasets, the approach achieves a 61.3% mean zero‑shot AUC, outperforming CLAP and Qwen2‑Audio, and reaches the highest linear probing AUC with only 43% of the data used by full‑scale baselines.
By Mustafa Talha \.Ilerisoy, Hung Manh Pham, Mathias Funk, Mykola Pechenizkiy, Aaqib Saeed
arXiv:2606. 28419v2 Announce Type: replace-cross Abstract: Limited data availability, class imbalance, and domain variability remain major barriers to reliable medical image classification.
By Teerath Kumar, Raja Vavekanand, Muhammad Turab
arXiv:2608.21300v1 Announce Type: new
Abstract: Foundation models for medical image segmentation, like prompt-based MedSAM, generalize well across domains and modalities, often in zero or few-shot se...
By Marko Haralovi\'c, Sounic Akkaraju, Carlo Baretta, Vasil Zapryanov, Alexia Briassouli
arXiv:2606. 28419v1 Announce Type: cross Abstract: Limited data availability, class imbalance, and domain variability remain major barriers to reliable medical image classification.
By Teerath Kumar, Raja Vavekanand, Muhammad Turab
arXiv:2606. 15436v1 Announce Type: cross Abstract: Respiratory acoustic foundation models (FMs) excel at cough classification, yet their ability to predict continuous health quantities from cough audio remains largely unexplored, despite the clinical value of passive age, BMI, and disease probability estimation in settings where physical measurements are unavailable.
By Mayur Sanap, Prasanna Desikan, Edgar Lobaton
The paper introduces the Cross‑Modal Triage Network (CMTN), a multimodal deep‑learning model that fuses a Swin Transformer V2 visual encoder with a PubMedBERT text encoder to perform severity‑based triage, pathology detection, and generate visual explanations for chest radiographs. Trained on 34,639 image‑text pairs from MIMIC‑CXR‑JPG, the CMTN achieves high ordinal agreement with reference labels (QWK = 0.9341) and excellent pathology detection (macro‑AUROC = 0.9970) while operating with 34 ms latency. However, a blinded clinical audit revealed low agreement with expert radiologists (QWK = 0.1399) and only modest spatial‑semantic concordance in heatmaps, underscoring the gap between algorithmic performance and clinical judgment.
By Zinah Ghulam, Richa Mittal, Eranga Ukwatta
arXiv:2608.30467v1 Announce Type: new
Abstract: Deep-learning models can achieve strong chest X-ray (CXR) classification performance without establishing whether their predictions predominantly rely...
By Abdullah Al Mamun, Md. Nasif Osman Khansur, Md Ashraful Hossen Akash, Md. Kishor Morol, Tze Hui Liew
arXiv:2608.23373v1 Announce Type: new
Abstract: Forecasting long-range influenza-like illness (ILI) matters for public health readiness. Publicly available surveillance datasets typically pair numeri...
By Seyed Mohammad Hossein Hashemi, Mohsen Hooshmand, Parvin Razzaghi