arXiv Computer Vision

SPECTRA: Band-Routed Embedding and Stage-Wise LoRA for Cross-Sensor Fine-Tuning of Geospatial Foundation Models

SPECTRA is a parameter‑efficient fine‑tuning framework for geospatial foundation models that tackles two key challenges: spectral mismatch and high adaptation cost. It introduces Band‑Routed Embedding (BRE) to map downstream sensor bands into the pretrained model’s expected band space, enabling full use of available spectral data without altering the patch embedding interface. Additionally, Stage‑wise Transferability‑aware LoRA (ST‑LoRA) estimates stage‑wise transferability and assigns LoRA ranks accordingly, concentrating trainable parameters on the most transferable stages and reducing overall adaptation cost.

arXiv Machine Learning
Jun 18

Task-Adaptive Parameter-Efficient Fine-Tuning for Weather Foundation Models

arXiv:2509. 22020v2 Announce Type: replace Abstract: While recent advances in machine learning have equipped Weather Foundation Models (WFMs) with substantial generalization capabilities across diverse downstream tasks, the escalating computational requirements associated with their expanding scale increasingly hinder practical deployment.

By Shilei Cao, Hehai Lin, Jiashun Cheng, Yang Liu, Guowen Li, Xuehe Wang, Juepeng Zheng, Haoyuan Liang, Meng Jin, Chengwei Qin, Hong Cheng, Haohuan Fu
arXiv Machine Learning
Sep 23

Parameter-Efficient Adaptation of Pre-Trained Vision Foundation Models for Active and Passive Seismic Data Denoising

The paper presents a framework that adapts general-purpose Vision Foundation Models (VFMs) to seismic data denoising using Parameter‑Efficient Fine‑Tuning with Low‑Rank Adaptation (LoRA). It introduces a kurtosis‑guided unsupervised test‑time adaptation module that updates only LoRA parameters to self‑calibrate for site‑specific noise without ground truth. Experiments on exploration seismic images and DAS data demonstrate that the approach matches or surpasses domain‑specific models and generalizes well to unseen cross‑site data.

By Jiahua Zhao, Umair bin Waheed, Jing Sun, Yang Cui, Nikos Savva, Eric Verschuur
arXiv Computer Vision
Aug 31

HyperVision: A Channel-Adaptive Ground-Based Hyperspectral Vision Pre-trained Backbone

HyperVision introduces the first ground‑based hyperspectral pre‑trained backbone, addressing challenges of varying spectral configurations, limited annotations, and dataset diversity. It employs a channel‑adaptive dynamic embedding to unify heterogeneous inputs, a multi‑source pseudo‑labeling strategy combining SAM2 spatial cues with HyperFree spectral details, and cross‑modal knowledge distillation from a pre‑trained RGB vision model. Trained on 15k images from 26 datasets, HyperVision achieves significant improvements—up to 16.3% relative gain in hyperspectral semantic segmentation, 2.1% in object tracking AUC, and 35.5% reduction in salient object detection MAE—while requiring only head‑only adaptation.

By Guanyiman Fu, Jingtao Li, Zihang Cheng, Zhuanfeng Li, Diqi Chen, Yan Xu, Xiangyu Liu, Fengchao Xiong, Jianfeng Lu, Chengrong Chen, Jun Zhou
arXiv AI
Aug 10

SLED: Scalable Location Encoding via Distillation

arXiv:2608. 06612v1 Announce Type: cross Abstract: The plethora of readily available geospatial data offers exciting opportunities to learn high quality representations of the planet, but the sheer size of the Earth Observations (EO), differing modalities, and different sensor types pose significant challenges in doing so.

By Kevin Lane, Zhongying Wang, Esther Rolf, Morteza Karimzadeh
arXiv Computer Vision
Aug 28

SIMPLER: Efficient Foundation Model Adaptation via Similarity-Guided Layer Pruning for Earth Observation

SIMPLER is a pre‑fine‑tuning method that reduces inference and deployment costs for Earth Observation foundation models by pruning redundant layers. It uses layer‑wise representation similarity on unlabeled task data to identify and remove up to 79% of parameters without requiring gradients, magnitude heuristics, or hyperparameter tuning. Experiments on Prithvi‑EO‑2, TerraMind, and ImageNet‑pretrained ViT‑MAE show that SIMPLER retains 94% of baseline performance while achieving 2.1× faster training and 2.6× faster inference.

By V\'ictor Barreiro, Johannes Jakubik, Francisco Arg\"uello, Dora B. Heras
arXiv AI
Jun 16

GeoRoPE: Ground-Aware Rotary Adaptation for Remote Sensing Foundation Models

arXiv:2606. 14760v1 Announce Type: cross Abstract: Remote-sensing foundation models (RSFMs) benefit from pretraining on imagery from multiple sensors and ground sampling distances (GSDs), but such exposure alone does not resolve scale mismatch during downstream adaptation.

By Yu Luo, Kun Hu, Mengwei He, Xiaogang Zhu, Shan Zeng, Allen Benter, Wei Xiang, Patrick Filippi, Thomas Francis Bishop, Zhiyong Wang