arXiv:2609.15888v1 Announce Type: cross
Abstract: Deep networks trained on structural MRI for Alzheimer's disease (AD) staging often reach reasonable accuracy while attending to anatomically irreleva...
By Paul-Gabriel Nicolae, Irina Georgiana Mocanu
The study investigates whether a compact, supervised 3D CNN pretrained for brain‑age prediction can act as a reusable foundation model for various Alzheimer's‑related neuroimaging tasks. By freezing the 7.18 million weights and adding only ~1 % of trainable parameters via Low‑Rank Adaptation, the model achieved high performance across six experiments, including dementia classification, MCI progression prediction, amyloid positivity detection, and volume estimation of hippocampal and white matter hypointensities. The results demonstrate that the pretrained brain‑age model generalizes well to new datasets without retraining, offering a data‑efficient alternative to larger networks.
By Reza Rajabli, D. Louis Collins
The paper introduces a lightweight Cross‑Layer Fusion Adapter (CLFA) that adapts the CLIP vision‑language model for handwriting‑based Alzheimer's disease screening. CLFA inserts multi‑level adapters into a frozen visual encoder, fusing cross‑layer features with depthwise 2D convolutions to capture both local stroke irregularities and higher‑level handwriting structure. On the Darwin dataset, CLFA achieves 74.63% AUC, 74.85% accuracy, and 73.72% F1, outperforming the best competing model by 2.15, 1.79, and 1.87 percentage points across 600 task‑disjoint source‑target pairs.
By Changqing Gong, Huafeng Qin, Mounim A. El-Yacoubi
arXiv:2606. 06534v1 Announce Type: cross Abstract: Longitudinal medical visual question answering (VQA) requires reasoning about anatomical differences between an image of a current time point and an image of a referred time point.
By Jialin Wu, Qianru Zhang, Georges El Fakhri, Xiaofeng Liu
The paper introduces Bidirectional Reciprocal Learning (BRL), a parameter‑efficient fine‑tuning framework for referring image segmentation that operates on frozen vision foundation models. BRL employs two lightweight adapters—Reciprocal Attention Adapter (RAA) for token‑level cross‑modal attention and Reciprocal Gate Adapter (RGA) for channel‑level gating—to enable hierarchical, bidirectional information flow between vision and language. Experiments on RefCOCO, RefCOCO+, and RefCOCOg show that BRL outperforms existing methods while updating fewer than 0.5% of backbone parameters.
By Xiaoqiang Lu, Licheng Jiao, Lingling Li, Yuting Yang, Long Sun, Wenping Ma, Xu Liu, Fang Liu
Rhamba is a region‑aware pretraining framework for resting‑state fMRI that combines anatomically guided masking with hybrid Attention‑Mamba architectures. The study pretrained models on the ABIDE dataset using three masking strategies (Any, Majority, Pure) and evaluated four architectural variants, finding that the Mamba‑Attention (MA) hybrid achieved the best average AUROC on downstream schizophrenia and ADHD classification tasks. Explainable AI via Integrated Gradients highlighted that performance depends on the interaction between masking strategy and architecture rather than a single dominant configuration.
By Pankaj Pandey, Ruthwik Reddy Doodipala, Pratheek Eranki, Carolina Torres-Rojas, Manob Jyoti Saikia, Ranganatha Sitaram
arXiv:2609.24510v2 Announce Type: replace
Abstract: Recent advances in vision foundation models (VFMs) have shown remarkable capabilities across diverse unimodal visual tasks. However, adapting VFMs...
By Xiaoqiang Lu, Licheng Jiao, Lingling Li, Yuting Yang, Long Sun, Wenping Ma, Xu Liu, Fang Liu
arXiv:2605. 18419v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs) can couple visual perception with open-ended clinical reasoning, making them attractive for computational histopathology.
By Franciskus Xaverius Erick, Johanna Paula M\"uller, Bernhard Kainz
arXiv:2606. 20037v1 Announce Type: new Abstract: Alzheimer's disease (AD) is an irreversible neurodegenerative disorder and a leading cause of death worldwide.
By Loukas Ilias, Anthi-Maria Vozinaki, Christos Ntanos, Dimitris Askounis
M$^2$PFN is an end‑to‑end multimodal framework that extends the TabPFN in‑context learning engine to Alzheimer’s disease diagnosis by aligning 3D‑MRI and tabular features in a shared subspace. It performs differentiable inference through TabPFN’s transformer, back‑propagates gradients into the encoders, and incorporates a frozen tabular‑only prediction via a gated shortcut. On the ADNI cohort it achieves 65.55 % macro‑F1 and 82.21 % macro‑AUC, surpassing unimodal and multimodal baselines, and it generalizes to external cohorts without retraining.
By Lujia Zhong, Shuo Huang, Jianwei Zhang, Xinyu Nie, Yonggang Shi
arXiv:2606. 04772v1 Announce Type: cross Abstract: Understanding the relationship between deep visual representations and the human visual system is a fundamental challenge in computational neuroscience.
By Hoang-Son Vo, Van-Hung Bui, Minh-Huy Mai-Duc, Tien-Dung Mai, Soo-Hyung Kim
arXiv:2608. 07749v1 Announce Type: cross Abstract: Parameter-efficient fine-tuning enables the adaptation of vision foundation models to biomedical tasks under limited computational resources, but a single low-rank update can constrain all task-specific changes to one narrow parameter subspace.
By Saed Moradi, Benyamin Ghojogh, M. Hadi Sepanj, Yimin Yang, Ashirbani Saha