arXiv AI

Zero-Shot Respiratory Sound Classification through LLM-Augmented Audio-Text Alignment

The paper introduces a framework that aligns self‑supervised respiratory encoders with medical terminology in a shared latent space, enabling zero‑shot respiratory sound classification. By using a medical LLM to generate structured reports from metadata, the method creates dense semantic anchors for contrastive learning, combining a sigmoid‑based contrastive loss with the encoder’s native SSL objective and similarity‑aware negative sampling. On nine tasks across six datasets, the approach achieves a 61.3% mean zero‑shot AUC, outperforming CLAP and Qwen2‑Audio, and reaches the highest linear probing AUC with only 43% of the data used by full‑scale baselines.

arXiv AI
Sep 18

FOCAL: Fine-Grained Optimal-Transport-Driven Contrastive Alignment of Language and ECGs with Waveform Enhancement

FOCAL is a framework that aligns fine-grained ECG waveform segments with specific report tags using Optimal Transport, addressing the lack of localized representation in prior methods. It introduces a semantic similarity matrix to mitigate false negatives when reports share diagnoses, and a coarse‑to‑fine enrichment pipeline that employs Large Language Models to recover missing waveform semantics while filtering hallucinations. Experiments on six datasets show FOCAL achieves state‑of‑the‑art zero‑shot prediction and linear probing performance.

By Haitao Li, Che Liu, Zhengyao Ding, Ziyi Liu, Wenqi Shao, Zhengxing Huang
Hugging Face Trending Papers
Jun 9

Closing the Modality Gap in Zero-Shot HAR: Contrastive Training and Separability-Optimized Prototypes on IMU Data

Zero-shot learning (ZSL) for inertial measurement unit (IMU)-based human activity recognition (HAR) faces a central challenge: bridging the gap between sensor embeddings and semantic class representations. We systematically evaluate seven configurations combining three inference methods with two training pipelines on the PAMAP2 dataset, using 14 seen and 4 unseen activity classes with subjects 108 and 109 held out for testing.

arXiv Computer Vision
Sep 3

AlphaRAD: Grounded Zero-Shot Classification in Chest Radiology via $\alpha$-Corrected Binary Cross Entropy and Factorized Latent Supervision

AlphaRAD introduces a grounded zero‑shot classification framework for chest radiology that leverages structured medical concepts extracted from reports and a novel α‑Corrected Binary Cross‑Entropy loss to reduce in‑batch noise. It also presents FLaS, a lightweight cross‑modal fusion module that factorizes VLPM representations into independent subspaces, improving spatial grounding without adding parameters. The method achieves state‑of‑the‑art performance on 16 classification benchmarks and sets new records on several grounding, phrase‑grounding, and segmentation datasets.

By Jianzhong You, Yuan Gao, Chris McIntosh
arXiv Computation and Language
Sep 22

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving

arXiv:2607.01733v2 Announce Type: replace Abstract: Speech-LLM integration has shown promising results by leveraging extensive textual pretraining, yet its specific benefits for automatic speech reco...

By Ruchao Fan, Yiming Wang, Rui Zhao, Liliang Ren, Keqi Deng, Xiaoyang Chen, Ali Zare, Bo Ren, Yuxuan Hu, Junkun Chen, Yan Huang, Yelong Shen, Jinyu Li