arXiv:2609.39266v1 Announce Type: new
Abstract: Fine-grained vision-language alignment in chest radiography enables zero-shot classification, grounding, and segmentation without task-specific annotat...
By Qixing Zhao, Jinpeng Li
NeoRed is a multimodal large language model specifically designed for diagnosing neonatal respiratory diseases. It addresses two main limitations of existing models—domain gaps from adult data and inadequate integration of clinical context—by leveraging two real-world neonatal datasets (NeoCXR and NeoCXR-EV). The model incorporates a Knowledge-Logic-Alignment framework that injects diagnostic priors, aligns report semantics with diagnostic logic, and aligns visual features with imaging conclusions, achieving superior performance on neonatal benchmarks while maintaining adult benchmark performance.
By Yinan Liu, Hongtai Xia, Haoran Xu, Jiankang Hong, Jingkuan Song, Ye Luo
arXiv:2606. 16484v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) hold great potential for medicine, as they inherit knowledge from LLM and allow multiple data modalities to be integrated, analysed and interpreted in natural language.
By Zhiyun Song, Che Liu, Tian Xia, Avinash Kori, Wenjia Bai
UniAR is a unified framework that improves autism spectrum disorder (ASD) recognition by using multi-granularity prompt learning and a large multimodal model to generate diagnostic descriptions at word, phrase, and sentence levels. It aligns these semantic representations with visual evidence through a Mixture-of-Experts-based Multi-Scale Alignment Module, enabling robust ASD detection across heterogeneous data types. Experiments on four brain MRI and facial expression benchmarks show that UniAR outperforms state‑of‑the‑art methods, achieving 75.9% accuracy on MRI and 91.6% on facial benchmarks, with gains of 1.5 and 1.2 percentage points respectively.
By Lei Xin, Zeheng Wang, Jiayin Zhu, Shihong Huang, Fanhu Zeng, Changjiang Jiang, Dengbo He, Yutao Yue, Zhenglun Kong
arXiv:2608. 12689v1 Announce Type: cross Abstract: Multi-parametric magnetic resonance imaging (mpMRI) is a cornerstone for brain tumor diagnosis and treatment, yet current AI models face critical limitations: their lack of natural language interaction and interpretability impedes spatial information integration and cross-modal reasoning required clinically.
By Zhi Qiao, Xintong Wu, Yichu He, Feng Shi
arXiv:2409.16183v2 Announce Type: replace
Abstract: Radiology is a vital and complex component of modern clinical workflow and covers many tasks. Recently, vision-language (VL) foundation models in m...
By Xiaohong Liu, Guoxing Yang, Yulin Luo, Jiaji Mao, Xiang Zhang, Haibo Wang, Zhiyang He, Ming Gao, Shanghang Zhang, Jun Shen, Guangyu Wang
arXiv:2604. 16878v2 Announce Type: replace Abstract: Early prediction of severe clinical deterioration and remaining length of stay can enable timely intervention and better resource allocation in high-acuity settings such as the ICU.
By Zhongyuan Liang, Junhyung Jo, Hyang-Jung Lee, Sang Kyu Kim, Irene Y. Chen
The paper introduces Neuro‑JEPA, a sparse multimodal foundation model that learns unified representations of brain MRI across T1w, T2w, and FLAIR sequences using a latent predictive objective and a Mixture‑of‑Experts architecture. It was pretrained on over 1.5 million scans from 428,647 studies and systematically evaluates architectural, masking, objective, and sparsity choices for robust multimodal representation learning. Across 47 tasks from three health systems and 12 public datasets, Neuro‑JEPA consistently outperforms a simple CNN baseline, demonstrating its effectiveness for both clinical and research applications.
By Haoxu Huang, Long Chen, Jingyun Chen, Jinu Hyun, James Ryan Loftus, Kara Melmed, Daniel Orringer, Jennifer Frontera, Seena Dehkharghani, Arjun Masurkar, Narges Razavian
arXiv:2606. 17115v1 Announce Type: cross Abstract: Foundation models (FMs) have emerged as powerful representation extractors for medical data, yet their generalizability to datasets under distribution shift remains underexplored.
By Jingyu Hu, Giuseppe Tripodi, Reed Naidoo, Sarah F. McGough, Tapabrata Chakraborti
arXiv:2606. 06696v1 Announce Type: cross Abstract: Vision and language models (VLMs) hold immense promise to transform biomedical imaging workflows, from detecting lesions in chest X-rays to profiling cellular features in microscopy.
By Ryan D'Cunha, Alejandro Lozano, Xiaoxiao Sun, Daniel Vela Jarquin, Min Woo Sun, Josiah Aklilu, James Burgess, Yuhui Zhang, Ryan Nayebi, Paola Avila, Robayo, Jin Ye, Ming Hu, Zhongying Deng, Junjun He, Xin Chen, Yue Yao, Robert Tibshirani, Jeffrey J. Nirschl, Serena Yeung-Levy
arXiv:2606. 17989v1 Announce Type: cross Abstract: Multi-contrast magnetic resonance imaging (MRI) provides complementary information for clinical diagnosis.
By Yonghao Chen, Sicheng Yang, Rui Tang, Lei Zhu
SCOUT is a concept‑grounded multimodal transformer that generates whole‑slide pathology reports by integrating local histological patterns, whole‑slide context, and expert‑curated diagnostic concepts. It uses evolving visual representations and recursively updated slide‑ and concept‑conditioned representations, with separate attention pathways during decoding that are fused adaptively for each token. Evaluated on TCGA‑BRCA, HistAI, and REG‑2025, SCOUT outperformed existing methods, improving BLEU, METEOR, and ROUGE‑L scores and raising the Clinical Report Quality Score on REG‑2025.
By Suryakant Singh, Saarthak Kapse, Joel Saltz, Prateek Prasanna