arXiv:2608. 04472v1 Announce Type: cross Abstract: The development of foundation models (FMs) is crucial for advancing endoscopic image analysis.
By Zhenyu Yi, Jianwei Xu, Yue Hu, Zhongwei Qiu, Sijing Li, Liang Huang, Bin Lv, Ling Zhang, Yingda Xia
arXiv:2608. 19825v1 Announce Type: cross Abstract: Medical image captioning is a technique that accelerates early-stage diagnostic workflows and enhances the interpretability of medical diagnostic AI systems.
By Yunseo Lee, Hyun Jun Kim, Heeseung Shin, Changwon Lim
arXiv:2609.14467v1 Announce Type: cross
Abstract: Automating clinical documentation from long-form doctor-patient conversations remains challenging for modern audio-language models. While cascaded AS...
By Ziyu Zhang, Mingchen Shao, Wenjie Tian, Tianlun Zuo, Longhao Li, Lei Xie
arXiv:2606. 17339v1 Announce Type: new Abstract: Speech offers a uniquely informative window into health by simultaneously engaging neurological, motor, respiratory, and vocal systems.
By Sejal Bhalla, Larry Kieu, Aina Merchant, Eyal de Lara, Alex Mariakakis
arXiv:2606. 10789v1 Announce Type: new Abstract: Zero-shot learning (ZSL) for inertial measurement unit (IMU)-based human activity recognition (HAR) faces a central challenge: bridging the gap between sensor embeddings and semantic class representations.
By Anik Ghosh
FOCAL is a framework that aligns fine-grained ECG waveform segments with specific report tags using Optimal Transport, addressing the lack of localized representation in prior methods. It introduces a semantic similarity matrix to mitigate false negatives when reports share diagnoses, and a coarse‑to‑fine enrichment pipeline that employs Large Language Models to recover missing waveform semantics while filtering hallucinations. Experiments on six datasets show FOCAL achieves state‑of‑the‑art zero‑shot prediction and linear probing performance.
By Haitao Li, Che Liu, Zhengyao Ding, Ziyi Liu, Wenqi Shao, Zhengxing Huang