arXiv:2603.28387v3 Announce Type: replace-cross
Abstract: Trustworthy clinical AI must use real evidence and avoid relying on surface-level artifacts. We evaluate 12 open-weight vision-language model...
By Doan Nam Long Vu, Simone Balloccu
arXiv:2609.23983v1 Announce Type: new
Abstract: Neuroimaging foundation models pretrained on large, predominantly western cohorts are increasingly proposed as general-purpose backbones for brain MRI...
By Oluwatobi Iyanuoluwa Akinmuleya, Olatokun Shamsudeen Akano, Samuel Danquah Ankapong, Olamide Lawal, Toufiq Musah
The paper introduces Neuro‑JEPA, a sparse multimodal foundation model that learns unified representations of brain MRI across T1w, T2w, and FLAIR sequences using a latent predictive objective and a Mixture‑of‑Experts architecture. It was pretrained on over 1.5 million scans from 428,647 studies and systematically evaluates architectural, masking, objective, and sparsity choices for robust multimodal representation learning. Across 47 tasks from three health systems and 12 public datasets, Neuro‑JEPA consistently outperforms a simple CNN baseline, demonstrating its effectiveness for both clinical and research applications.
By Haoxu Huang, Long Chen, Jingyun Chen, Jinu Hyun, James Ryan Loftus, Kara Melmed, Daniel Orringer, Jennifer Frontera, Seena Dehkharghani, Arjun Masurkar, Narges Razavian
The paper introduces MedREAL, a unified framework that aligns linguistic reasoning with spatial grounding for medical visual question answering and segmentation. MedREAL employs Seg Anchored Reasoning Pooling (SARP) to extract semantic evidence from segmentation tokens and a Reasoning-to-Visual (R2V) fusion mechanism to integrate these features into a segmentation pipeline. Using the newly created MedRAVS-13K dataset, MedREAL achieves superior performance, reporting 68.49% gIoU and 70.47% cIoU, and generates evidence masks that consistently match textual diagnoses.
By Haowen Gu, Gensheng Pei, Junzhu Mao, Qiong Wang, Mingwu Ren, Yazhou Yao
The study presents an explainable multimodal deep‑learning framework that combines a 3D CNN for T1‑weighted MRI with a feedforward network for harmonized clinical and demographic data to diagnose Alzheimer’s disease. Using 6,479 ADNI records and 1,703 OASIS‑3 records, the authors compare various model configurations on three‑way and pairwise diagnostic tasks, finding that performance and explanations vary by task, modality, fusion strategy, and cohort. SHAP and Integrated Gradients consistently highlight the MMSE score as the most influential tabular feature, while CAM‑based explanations differ across model setups and cohorts, indicating that explainability is not a stable property under cohort shift.
By Yusuf Brima, Marcellin Atemkeng, Lakshmana Rao Namamula, Antoine Vacavant
arXiv:2607. 16325v1 Announce Type: cross Abstract: Foundation models provide powerful representations for brain MRI analysis, but their predictions remain difficult to interpret in anatomically meaningful terms.
By Wei Zhang
arXiv:2608.22323v1 Announce Type: new
Abstract: The application of Large Language Models (LLMs) to diagnostic decision-making has garnered growing interest. However, existing benchmarks largely focus...
By Lai Wei, Yuchao Chen, Zhenbiao Cao, Xiaojin Zhang, Zhongyu Wei, Bangting Wang, Wei Chen, Xiang Bai
arXiv:2608. 12689v1 Announce Type: cross Abstract: Multi-parametric magnetic resonance imaging (mpMRI) is a cornerstone for brain tumor diagnosis and treatment, yet current AI models face critical limitations: their lack of natural language interaction and interpretability impedes spatial information integration and cross-modal reasoning required clinically.
By Zhi Qiao, Xintong Wu, Yichu He, Feng Shi
arXiv:2609.00593v1 Announce Type: new
Abstract: Neuroradiologists rarely read a brain MRI in isolation, yet automated brain-MRI report generation has been built almost entirely for single studies. Te...
By Krish Patel, Peirong Liu
The paper investigates why multimodal large language models (MLLMs) struggle with vision‑centric tasks when visual evidence conflicts with pretrained language knowledge. Using image reconstruction and a new WhatIfVis benchmark, the authors show that MLLMs preserve coarse‑grained visual attributes but fail to consistently use them, and that supervised fine‑tuning and activation patching can improve controllability of visual context sensitivity. The study demonstrates that the main bottleneck lies in the models’ inability to reliably regulate their reliance on visual evidence rather than in visual perception itself.
By Jiaang Li, Chengzu Li, Zhaochong An, Yifei Yuan, Xi Liu, Serge Belongie, V\'esteinn Sn{\ae}bjarnarson
arXiv:2607. 16303v1 Announce Type: cross Abstract: Medical Vision-Language Models (Med-VLMs) require reliable reasoning from fine-grained visual evidence, yet existing models can produce plausible clinical answers by relying on language priors or medical templates rather than truly attending to diagnosis-critical regions.
By Yunhang Qian, Jiaquan Yu, Jiawei Liu, Meng Wang, Hongwei Bran Li, Xiaobin Hu
arXiv:2609.15888v1 Announce Type: cross
Abstract: Deep networks trained on structural MRI for Alzheimer's disease (AD) staging often reach reasonable accuracy while attending to anatomically irreleva...
By Paul-Gabriel Nicolae, Irina Georgiana Mocanu