arXiv Machine Learning

When Bigger is Worse: A Practitioner's Guide to Model Selection Under Data Scarcity

arXiv:2603. 02142v2 Announce Type: replace-cross Abstract: Scaling laws assume larger models trained on more data consistently outperform smaller ones -- an assumption that drives model selection in computer vision but remains untested in resource-constrained Earth observation (EO).

arXiv Machine Learning
Jul 7

TESSERA v2: Scaling Pixel-wise Earth Foundation Models

arXiv:2607. 03949v1 Announce Type: cross Abstract: Pixel-wise Earth-observation (EO) foundation models are now achieving state-of-the-art performance via generated spatial embeddings.

By Zhengpeng Feng, Sadiq Jaffer, Ira Shokar, Jovana Knezevic, Mark Elvers, Clement Atzberger, Robin Young, Aneesh Naik, Niall Robinson, Andrew Blake, David Coomes, Anil Madhavapeddy, Srinivasan Keshav
arXiv Computer Vision
Aug 28

SIMPLER: Efficient Foundation Model Adaptation via Similarity-Guided Layer Pruning for Earth Observation

SIMPLER is a pre‑fine‑tuning method that reduces inference and deployment costs for Earth Observation foundation models by pruning redundant layers. It uses layer‑wise representation similarity on unlabeled task data to identify and remove up to 79% of parameters without requiring gradients, magnitude heuristics, or hyperparameter tuning. Experiments on Prithvi‑EO‑2, TerraMind, and ImageNet‑pretrained ViT‑MAE show that SIMPLER retains 94% of baseline performance while achieving 2.1× faster training and 2.6× faster inference.

By V\'ictor Barreiro, Johannes Jakubik, Francisco Arg\"uello, Dora B. Heras
arXiv AI
2d ago

Useful to Whom? Sample Value Is Defined Only Relative to the Learner

The paper investigates how the usefulness of training samples, as determined by coreset selection, depends on the learner rather than just the data. Experiments on ImageNet-100 and ImageNet-1k show that changing model width, input grid, stride, and architecture (e.g., ResNet vs. ViT) shifts the crossover point where different selection criteria (easy-first vs. geometric coverage) become optimal. These findings demonstrate that the relative value of a fixed subset of samples varies with the target learner’s capacity and structure, and that selection strategies must be tuned to the specific model they will train.

By Yangze Liu, Xiao-Long Yin, Zhongyi Han
arXiv Computer Vision
4d ago

RS-OPSD: Reliable Privileged On-Policy-Self-Distillation for Ultra-High-Resolution Remote Sensing VQA

The paper introduces RS-OPSD, a reliable privileged on-policy self-distillation framework designed for ultra‑high‑resolution remote sensing visual question answering. It leverages a new dataset, GeoEvidence‑6K, and a human‑feedback guided skill refinement process to provide explicit question‑relevant evidence. By incorporating context‑preserving visual privilege and correctness‑aligned distillation, RS‑OPSD achieves state‑of‑the‑art performance on several benchmarks without requiring additional visual search or tool calls at inference time.

By Chengjie Jiang, Yunqi Zhou, Jiafeng Yan, Sihang Zhao, Chun Yuan, Jing Li
arXiv Computer Vision
3d ago

Cryo-Bench: Benchmarking Foundation Models for Cryosphere Mapping

Cryo-Bench is a new benchmark that evaluates foundation models for cryosphere mapping, comprising six semantic‑segmentation datasets across five cryospheric components (supraglacial debris, glacial lakes, sea ice, calving fronts, and Antarctic ice‑shelf extent). The benchmark includes multispectral, RGB, and SAR observations from under‑represented regions and tests thirteen geo‑foundation models alongside U‑Net and Vision Transformer baselines. Results show that with frozen encoders U‑Net slightly outperforms TerraMind, but the difference is not statistically significant; fine‑tuning with learning‑rate optimization can dramatically improve performance for some models, while in few‑shot scenarios several foundation models retain over 90 % of their full‑label accuracy.

By Saurabh Kaushik, Lalit Maurya, Beth Tellman, Swalpa Kumar Roy, Valerio Marsocci, Gustau Camps-Valls, Jocelyn Chanussot
arXiv AI
Aug 28

Successive Capacity Growth: Task-Complexity-Driven Width and Depth Expansion for Vision Transformer Encoders in JEPA World Models

The paper introduces Successive Capacity Growth (SCG), a method for adaptively expanding Vision Transformer encoders in Joint-Embedding Predictive Architectures (JEPAs). SCG starts with a minimal encoder and incrementally increases width or depth based on a task‑agnostic test‑and‑verify mechanism, while a Sketched Isotropic Gaussian Regularizer (SIGReg) keeps learned semantic dimensions independent. Experiments on multi‑object dynamics and 2D navigation tasks show that SCG achieves up to 20.3% better prediction loss than fixed small baselines and 23% better than fixed large models, with far greater parameter efficiency and no false‑positive expansions.

By Frederik Berenz
arXiv AI
Aug 25

What Transfers from Text to Vision? Capability Scaling Laws and Transfer Dynamics for VLMs

The paper introduces the Capability-Driven Multimodal Scaling Law, a cross-family framework that predicts vision-language model (VLM) benchmark accuracy from a low-dimensional textual capability score extracted via PCA. By training over 150 VLMs on 34 large language models across seven families, the authors demonstrate that the law accurately extrapolates transfer rates from 8B to 72B‑parameter backbones, predicts full training trajectories, and generalizes to unseen model families. The study also reveals actionable insights, such as certain textual benchmarks negatively correlating with multimodal performance and base LLMs outperforming instruction-tuned counterparts as VLM backbones due to higher absorption rates.

By Ziran Li, Qiang Wang, Zhengyu Chen, Shanglin Lei, Borun Chen, Jingang Wang, Xunliang Cai
Hugging Face Trending Papers
Aug 27

Successive Capacity Growth: Task-Complexity-Driven Width and Depth Expansion for Vision Transformer Encoders in JEPA World Models

The paper introduces Successive Capacity Growth (SCG), a method that starts with a minimal Vision Transformer encoder and incrementally expands its width or depth based on a task‑agnostic test‑and‑verify mechanism. SCG uses function‑preserving expansion and a Sketched Isotropic Gaussian Regularizer (SIGReg) to ensure independent semantic dimensions and prevent collapse. Experiments on multi‑object dynamics and 2D navigation tasks show that SCG achieves significant prediction loss reductions while being far more parameter‑efficient than fixed large models, with no false‑positive expansions and exact function preservation.

arXiv Computer Vision
Sep 23

Vision Foundation Models with Synthetic-Only Training for Monocular Spacecraft Pose Estimation

The paper presents a new monocular spacecraft pose estimation model that achieves the lowest reported mean rotation errors on the SPEED+ lightbox and sunlamp test sets. By replacing smaller encoders with a large self‑supervised ViT foundation model (DINOv3) and scaling up to 840 M parameters, the authors improve accuracy from 300 M to 840 M parameters without saturation. The 840 M model also runs on a Jetson Orin NX 16 GB with 133.8 ms per crop and 32.0 W power draw, demonstrating embedded inference feasibility while training solely on synthetic data.

By John Church, Vazghen Nikolian