arXiv Computer Vision

FedSSMCoOp: SSM Encoders for light-weight Federated Prompt Learning for Few-shot Classification

arXiv Machine Learning
Jul 20

Toward Federated Multimodal Graph Foundation Models: A Topology-Aware Multimodal Alignment Framework

arXiv:2607. 15687v1 Announce Type: new Abstract: Multimodal-attributed graphs (MAGs), whose nodes carry modalities such as images and text alongside topological structure, now pervade applications including social platforms, e-commerce, and biomedical networks, offering richer semantic signals than single-modality graphs.

By Xunkai Li, Guohao Fu, Yuming Ai, Zhengyu Wu, Hongchao Qin, Rong-Hua Li, Guoren Wang
arXiv AI
Sep 3

Federated LoRA Adaptation of BiomedCLIP Across Four International Chest X-Ray Cohorts

The paper evaluates federated learning with Low‑Rank Adaptation (LoRA) for fine‑tuning the BiomedCLIP vision‑language model on chest X‑ray classification across four international cohorts. Federated LoRA improves shared‑class AUC from 0.687 to 0.802, outperforming isolated single‑cohort training and approaching a centralized reference. The study shows that SVD‑based product‑space aggregation (FlexLoRA) is crucial for performance, while FedProx offers no advantage over FedAvg in this setting.

By Sanjaya Poudel, Nirajan Kunwor, Manish Dhakal, Debesh Jha, Sunil Kumar Gaire
arXiv Computer Vision
Sep 22

AdaptiveCDM: Source-Free Few-Shot Domain Adaptation for Cell Detection in Microscopic Images

AdaptiveCDM is a modular framework for source‑free few‑shot domain adaptation in cell detection, enabling a pretrained model to adapt to new imaging domains using only a handful of labeled target images and no source data. It combines Resolution‑Aware Augmentation (RAug) to balance scarce, class‑imbalanced samples while preserving cellular morphology, and Category‑Aware Representation Learning (CARL) to strengthen class‑consistent proposals for better localization and classification. Experiments on M5 and Raabin‑WBC datasets show that AdaptiveCDM achieves competitive or superior mAP scores compared to state‑of‑the‑art methods under their respective supervision settings.

By Nimra Dilawar, Sara Nadeem, Javed Iqbal, Waqas Sultani, Mohsen Ali
arXiv AI
Jun 8

MMBU: A Massive Multi-modal Biomedical Understanding Benchmark to Probe the Perception Capabilities of Vision-Language Models

arXiv:2606. 06696v1 Announce Type: cross Abstract: Vision and language models (VLMs) hold immense promise to transform biomedical imaging workflows, from detecting lesions in chest X-rays to profiling cellular features in microscopy.

By Ryan D'Cunha, Alejandro Lozano, Xiaoxiao Sun, Daniel Vela Jarquin, Min Woo Sun, Josiah Aklilu, James Burgess, Yuhui Zhang, Ryan Nayebi, Paola Avila, Robayo, Jin Ye, Ming Hu, Zhongying Deng, Junjun He, Xin Chen, Yue Yao, Robert Tibshirani, Jeffrey J. Nirschl, Serena Yeung-Levy
arXiv Machine Learning
Jul 9

MedPMC: A Systematic Framework for Scaling High-Fidelity Medical Multimodal Data for Foundation Models

arXiv:2607. 07673v1 Announce Type: cross Abstract: Medicine is inherently multimodal, requiring clinicians to synthesize information across diverse data streams.

By Hyunjae Kim, Dain Kim, Pan Xiao, Serina S. Applebaum, Younjoon Chung, Xuguang Ai, Yu Yin, Roy Jiang, Yuexi Du, Yawen Wei, Yiming Kong, Tuo Guo, Zhiyuan Cao, Mengmeng Du, Yuelei Fu, Yan Hu, Rui Shi, Gui Yang, Kevin W. Jin, Yuntian Liu, Yuxuan Tian, Jonathan Marquez, Zhen Chen, Sheng Zhang, Hoifung Poon, Hua Xu, Jaewoo Kang, Qingyu Chen
arXiv Computer Vision
Sep 28

ProCAP: Probabilistic Cross-Attentive Prompt Learning for Vision-Language Models

ProCAP introduces a probabilistic cross-attentive prompt learning framework for vision-language models like CLIP, enabling improved cross-modal interaction without updating the backbone. It jointly learns visual and textual prompt tokens, linking them via stacked bidirectional multi-head cross-attention to refine each branch across prompt depth. The method incorporates Gaussian parameterization of prompt tokens, lightweight KL and L2 regularization, and a compact symmetric InfoNCE head to align image features with class-level text representations, achieving strong few-shot base-to-novel performance and competitive transfer results across multiple datasets and benchmarks.

By Hiwa Azeez Abbas, Fatemeh Daneshfar, Moloud Abdar
arXiv Machine Learning
Sep 24

FFM-CP: Cross-Backbone Fusion of Vision-Language Foundation Models for Few-Shot Computational Pathology

The paper introduces FFM-CP, a framework that fuses multiple pathology vision‑language foundation models for few‑shot learning. It aligns heterogeneous representations with an Orthogonal Procrustes transformation, then uses a unified graph to refine support‑image features and class prototypes across backbones. Experiments on six histopathology datasets show that FFM‑CP outperforms the best single adapted model in 50 of 54 few‑shot comparisons.

By Anh-Tien Nguyen, Trung DQ. Dang, Nghiem Tuong Diep, Bui Ngoc Han Nguyen, Tan-Ha Mai, Miriam Cindy Maurer, Phuong Hoa Nguyen, Thi Thuy Uyen Nguyen, Youngjun Park, Daniel Sonntag, Duy Minh Ho Nguyen, Anne-Christin Hauschild
arXiv AI
Jun 24

Evaluating the Interpretability of Sparse Autoencoders with Concept Annotations

arXiv:2606. 24716v1 Announce Type: cross Abstract: Sparse autoencoders (SAEs) are increasingly used to extract interpretable concepts from vision and vision language models, yet existing evaluation methods largely rely on proxy metrics or qualitative inspection rather than measuring semantic correspondence.

By Jonas Klotz, Cassio F. Dantas, Pallavi Jain, Diego Marcos, Beg\"um Demir