arXiv Machine Learning By Sara Ketabi, Matthias W. Wagner, Cynthia Hawkins, Uri Tabori, Birgit Betina Ertl-Wagner, Farzad Khalvati

Multimodal Semantic-Aware Contrastive Learning For False Negative Mitigation in 3D Medical Imaging

Read the original on arXiv Machine Learning →

arXiv:2607. 14995v1 Announce Type: new Abstract: Multimodal Contrastive Learning (CL) has shown significant performance in aligning representations across various data modalities and improving downstream tasks, especially in healthcare.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 4

NeoRed: A Knowledge-Logic-Alignment Multimodal Large Language Model for Neonatal Respiratory Disease Diagnosis

NeoRed is a multimodal large language model specifically designed for diagnosing neonatal respiratory diseases. It addresses two main limitations of existing models—domain gaps from adult data and inadequate integration of clinical context—by leveraging two real-world neonatal datasets (NeoCXR and NeoCXR-EV). The model incorporates a Knowledge-Logic-Alignment framework that injects diagnostic priors, aligns report semantics with diagnostic logic, and aligns visual features with imaging conclusions, achieving superior performance on neonatal benchmarks while maintaining adult benchmark performance.

By Yinan Liu, Hongtai Xia, Haoran Xu, Jiankang Hong, Jingkuan Song, Ye Luo
arXiv AI
6d ago

UniAR: A Unified Framework for Autism Recognition Enhanced by Multi-View Prompt Learning

UniAR is a unified framework that improves autism spectrum disorder (ASD) recognition by using multi-granularity prompt learning and a large multimodal model to generate diagnostic descriptions at word, phrase, and sentence levels. It aligns these semantic representations with visual evidence through a Mixture-of-Experts-based Multi-Scale Alignment Module, enabling robust ASD detection across heterogeneous data types. Experiments on four brain MRI and facial expression benchmarks show that UniAR outperforms state‑of‑the‑art methods, achieving 75.9% accuracy on MRI and 91.6% on facial benchmarks, with gains of 1.5 and 1.2 percentage points respectively.

By Lei Xin, Zeheng Wang, Jiayin Zhu, Shihong Huang, Fanhu Zeng, Changjiang Jiang, Dengbo He, Yutao Yue, Zhenglun Kong
arXiv AI
Aug 14

Mr3D-VL: A generalist vision language foundation model for Multiparametric 3D Magnetic Resonance Imaging

arXiv:2608. 12689v1 Announce Type: cross Abstract: Multi-parametric magnetic resonance imaging (mpMRI) is a cornerstone for brain tumor diagnosis and treatment, yet current AI models face critical limitations: their lack of natural language interaction and interpretability impedes spatial information integration and cross-modal reasoning required clinically.

By Zhi Qiao, Xintong Wu, Yichu He, Feng Shi
arXiv Computer Vision
Aug 25

Expert-level vision-language foundation model for real-world radiology and comprehensive evaluation

arXiv:2409.16183v2 Announce Type: replace Abstract: Radiology is a vital and complex component of modern clinical workflow and covers many tasks. Recently, vision-language (VL) foundation models in m...

By Xiaohong Liu, Guoxing Yang, Yulin Luo, Jiaji Mao, Xiang Zhang, Haibo Wang, Zhiyang He, Ming Gao, Shanghang Zhang, Jun Shen, Guangyu Wang