arXiv AI By Doan Nam Long Vu, Simone Balloccu

The Scaffold Effect: How Prompt Framing Drives Apparent Multimodal Gains in Clinical VLM Evaluation

Read the original on arXiv AI →

arXiv:2603. 28387v2 Announce Type: replace Abstract: Trustworthy clinical AI requires that performance gains reflect genuine evidence integration rather than surface-level artifacts.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Sep 14

Learning Sparse Latent Predictive Foundation Model for Multimodal Neuroimaging

The paper introduces Neuro‑JEPA, a sparse multimodal foundation model that learns unified representations of brain MRI across T1w, T2w, and FLAIR sequences using a latent predictive objective and a Mixture‑of‑Experts architecture. It was pretrained on over 1.5 million scans from 428,647 studies and systematically evaluates architectural, masking, objective, and sparsity choices for robust multimodal representation learning. Across 47 tasks from three health systems and 12 public datasets, Neuro‑JEPA consistently outperforms a simple CNN baseline, demonstrating its effectiveness for both clinical and research applications.

By Haoxu Huang, Long Chen, Jingyun Chen, Jinu Hyun, James Ryan Loftus, Kara Melmed, Daniel Orringer, Jennifer Frontera, Seena Dehkharghani, Arjun Masurkar, Narges Razavian
arXiv AI
Aug 28

From Reasoning to Pixels: Grounded Medical Multimodal LLMs for VQA and Segmentation

The paper introduces MedREAL, a unified framework that aligns linguistic reasoning with spatial grounding for medical visual question answering and segmentation. MedREAL employs Seg Anchored Reasoning Pooling (SARP) to extract semantic evidence from segmentation tokens and a Reasoning-to-Visual (R2V) fusion mechanism to integrate these features into a segmentation pipeline. Using the newly created MedRAVS-13K dataset, MedREAL achieves superior performance, reporting 68.49% gIoU and 70.47% cIoU, and generates evidence masks that consistently match textual diagnoses.

By Haowen Gu, Gensheng Pei, Junzhu Mao, Qiong Wang, Mingwu Ren, Yazhou Yao
arXiv Computer Vision
Sep 14

A Multimodal Explainable Deep Learning Framework for Alzheimer's Disease Diagnosis using 3D Magnetic Resonance Imaging and Clinical Data

The study presents an explainable multimodal deep‑learning framework that combines a 3D CNN for T1‑weighted MRI with a feedforward network for harmonized clinical and demographic data to diagnose Alzheimer’s disease. Using 6,479 ADNI records and 1,703 OASIS‑3 records, the authors compare various model configurations on three‑way and pairwise diagnostic tasks, finding that performance and explanations vary by task, modality, fusion strategy, and cohort. SHAP and Integrated Gradients consistently highlight the MMSE score as the most influential tabular feature, while CAM‑based explanations differ across model setups and cohorts, indicating that explainability is not a stable property under cohort shift.

By Yusuf Brima, Marcellin Atemkeng, Lakshmana Rao Namamula, Antoine Vacavant