arXiv AI

Who Needs Labels? Adapting Vision Foundation Models With the Metadata You Already Have

arXiv:2606. 05107v1 Announce Type: cross Abstract: We propose a label-free approach to adapt powerful but generic vision foundation models to specialized scientific domains.

arXiv Computer Vision
Sep 22

AdaptiveCDM: Source-Free Few-Shot Domain Adaptation for Cell Detection in Microscopic Images

AdaptiveCDM is a modular framework for source‑free few‑shot domain adaptation in cell detection, enabling a pretrained model to adapt to new imaging domains using only a handful of labeled target images and no source data. It combines Resolution‑Aware Augmentation (RAug) to balance scarce, class‑imbalanced samples while preserving cellular morphology, and Category‑Aware Representation Learning (CARL) to strengthen class‑consistent proposals for better localization and classification. Experiments on M5 and Raabin‑WBC datasets show that AdaptiveCDM achieves competitive or superior mAP scores compared to state‑of‑the‑art methods under their respective supervision settings.

By Nimra Dilawar, Sara Nadeem, Javed Iqbal, Waqas Sultani, Mohsen Ali
arXiv Computer Vision
Aug 27

Semi-Supervised Adaptation of Vision-Language Models for Image Classification

The paper introduces Self‑Evolutionary CLIP (SE‑CLIP), a semi‑supervised framework that adapts vision‑language models like CLIP to satellite imagery. SE‑CLIP uses a two‑phase pipeline: an initial warm‑up on a small set of annotated seeds followed by a recursive discovery phase that iteratively selects high‑confidence samples from unlabeled data. A class‑balanced selection strategy is applied to keep the evolving support set balanced, and experiments on the UCM and NWPU benchmarks show that SE‑CLIP outperforms existing semi‑supervised methods.

By Mohamed L. Mekhalfi, Mohamad M. Al Rahhal, Yakoub Bazi, Salah E. Khenfer, Mingdeng Shi, Hua Zou, Mansour Zuair
arXiv Computer Vision
Sep 22

Toward a foundation model for forest point clouds

arXiv:2609.24787v1 Announce Type: new Abstract: Forest inventories increasingly rely on artificial intelligence (AI) models to derive forest attributes from large-scale 3D point clouds. Current model...

By Yuanwen Yue, Stefano Puliti, Damien Robert, Atakan Topalo\u{g}lu, Binbin Xiang, Maciej Wielgosz, Jan Dirk Wegner, Rasmus Astrup, Christian Rupprecht, Konrad Schindler
arXiv Machine Learning
Aug 4

OSMDA: OpenStreetMap-based Domain Adaptation for Remote Sensing VLMs

arXiv:2603. 11804v3 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) adapted to remote sensing rely heavily on domain-specific image-text supervision, yet high-quality annotations for satellite and aerial imagery remain scarce and expensive to produce.

By Stefan Maria Ailuro (INSAIT, Sofia University "St. Kliment Ohridski"), Mario Markov (INSAIT, Sofia University "St. Kliment Ohridski"), Mohammad Mahdi (INSAIT, Sofia University "St. Kliment Ohridski"), Delyan Boychev (INSAIT, Sofia University "St. Kliment Ohridski"), Luc Van Gool (INSAIT, Sofia University "St. Kliment Ohridski"), Danda Pani Paudel (INSAIT, Sofia University "St. Kliment Ohridski")
arXiv AI
Sep 25

A Multimodal 3D Foundation Model for Light Sheet Fluorescence Microscopy Enables Few-Shot Segmentation, Classification, and Deblurring

The paper presents a 3D foundation model for light sheet fluorescence microscopy (LSM) that is pretrained on a large curated set of 3D images from various organisms, stains, and imaging protocols. By jointly optimizing for masked reconstruction and image‑text alignment, the model learns transferable volumetric representations that dramatically reduce the need for annotated data. The pretrained backbone enables efficient few‑shot adaptation to downstream tasks such as segmentation, classification, and deblurring, consistently outperforming baselines according to standard metrics and expert evaluation.

By Adina Scheinfeld, Haotan Zhang, Shang Mu, Rudolf L. M. van Herten, Lucas Stoffl, Ali Erturk, Zhuhao Wu, Johannes C. Paetzold
arXiv Computer Vision
Aug 25

Retrieval-Augmented Visual Prompting: Guiding Foundation Models in Two-Photon Imaging

arXiv:2608.21970v1 Announce Type: new Abstract: Two-photon calcium imaging presents a challenging setting for foundation models: image appearance varies substantially across recordings and experiment...

By Salvatore Calcagno, Marco Finocchiaro, Giovanni Bellitto, Daniela Giordano, Concetto Spampinato, Federica Proietto Salanitri
arXiv AI
Aug 11

LoRSA: Toward Generalizable Parameter-Efficient Fine-Tuning for Biomedical Downstream Tasks

arXiv:2608. 07749v1 Announce Type: cross Abstract: Parameter-efficient fine-tuning enables the adaptation of vision foundation models to biomedical tasks under limited computational resources, but a single low-rank update can constrain all task-specific changes to one narrow parameter subspace.

By Saed Moradi, Benyamin Ghojogh, M. Hadi Sepanj, Yimin Yang, Ashirbani Saha