arXiv AI By Mustafa Talha \.Ilerisoy, Hung Manh Pham, Mathias Funk, Mykola Pechenizkiy, Aaqib Saeed

Zero-Shot Respiratory Sound Classification through LLM-Augmented Audio-Text Alignment

Read the original on arXiv AI →

The paper introduces a framework that aligns self‑supervised respiratory encoders with medical terminology in a shared latent space, enabling zero‑shot respiratory sound classification. By using a medical LLM to generate structured reports from metadata, the method creates dense semantic anchors for contrastive learning, combining a sigmoid‑based contrastive loss with the encoder’s native SSL objective and similarity‑aware negative sampling. On nine tasks across six datasets, the approach achieves a 61.3% mean zero‑shot AUC, outperforming CLAP and Qwen2‑Audio, and reaches the highest linear probing AUC with only 43% of the data used by full‑scale baselines.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 18

FOCAL: Fine-Grained Optimal-Transport-Driven Contrastive Alignment of Language and ECGs with Waveform Enhancement

FOCAL is a framework that aligns fine-grained ECG waveform segments with specific report tags using Optimal Transport, addressing the lack of localized representation in prior methods. It introduces a semantic similarity matrix to mitigate false negatives when reports share diagnoses, and a coarse‑to‑fine enrichment pipeline that employs Large Language Models to recover missing waveform semantics while filtering hallucinations. Experiments on six datasets show FOCAL achieves state‑of‑the‑art zero‑shot prediction and linear probing performance.

By Haitao Li, Che Liu, Zhengyao Ding, Ziyi Liu, Wenqi Shao, Zhengxing Huang