Large 3D foundation models such as MASt3R achieve state-of-the-art stereo reconstruction but are computationally demanding for deployment under strict hardware constraints -- a critical limitation in domains such as planetary exploration, where onboard computing is severely restricted. We study how far such models can be compressed through knowledge distillation, using lunar stereo reconstruction as a challenging and practically relevant case study.
The paper presents a method for distilling a large 300‑million‑parameter geospatial foundation model (Prithvi‑EO‑2.0) into a compact 0.7‑million‑parameter EfficientViT‑B0 student for flood segmentation. By using the teacher to supervise additional unlabeled Sentinel‑2 imagery, the student’s training set expands without new manual labels, achieving competitive performance on Sen1Floods11 and STURM‑Flood while remaining smaller and faster. After quantization, the student runs as a 1.5‑MB INT8 TensorRT engine on a Jetson Xavier NX, processing 512×512 images in 5.57 ms with ~14 MB of memory.
By Fabian Schmalstieg, Karsten Mueller, Wojciech Samek
arXiv:2609.38312v1 Announce Type: cross
Abstract: The Platonic Representation Hypothesis predicts that sufficiently scaled foundation models converge on a shared representation of the world. As each...
By Michael J. Smith, Shashwat Sourav
arXiv:2606. 00746v1 Announce Type: cross Abstract: Vision foundation models are bottlenecked by the quadratic cost of self-attention, which limits usable resolution and increases the cost of large-scale pretraining.
By Yitong Jiang, Hongjun Wang, Collin McCarthy, Hanrong Ye, David Wehr, Xinhao Li, Qi Dou, Tianfan Xue, Ka Chun Cheung, Simon See, Wonmin Byeon, Ke Chen, Kai Han, Jinwei Gu, Hongxu Yin, Pavlo Molchanov, Jan Kautz, Sifei Liu
arXiv:2603. 02142v2 Announce Type: replace-cross Abstract: Scaling laws assume larger models trained on more data consistently outperform smaller ones -- an assumption that drives model selection in computer vision but remains untested in resource-constrained Earth observation (EO).
By Kwame Mbobda-Kuate, Gabriel Kasmi
PixelDense introduces a dual‑stream representation alignment for pixel diffusion, separating semantic and geometric teachers (DINOv2, SAM2, Depth Anything v2, Metric3D v2) into distinct projection spaces with an orthogonality penalty. The method improves dense‑prediction benchmarks, boosting PixelGen‑XXL’s GenEval score from 0.7927 to 0.8093, achieving significant gains in panoptic quality and depth accuracy, and accelerating training from random initialization. It also enhances SDEdit editing by preserving background structure and increasing PSNR.
By Lehan Yang, Daiqing Qi, Wenhao Zhang, Avery Li, Yiqing Yang, Yifan Li, Yu Kong, Haitian Zheng, Zhifei Zhang, Zhe Lin, Varun Jampani, Sheng Li
SIMPLER is a pre‑fine‑tuning method that reduces inference and deployment costs for Earth Observation foundation models by pruning redundant layers. It uses layer‑wise representation similarity on unlabeled task data to identify and remove up to 79% of parameters without requiring gradients, magnitude heuristics, or hyperparameter tuning. Experiments on Prithvi‑EO‑2, TerraMind, and ImageNet‑pretrained ViT‑MAE show that SIMPLER retains 94% of baseline performance while achieving 2.1× faster training and 2.6× faster inference.
By V\'ictor Barreiro, Johannes Jakubik, Francisco Arg\"uello, Dora B. Heras
The study evaluates how the length of observation windows affects the performance of Tessera embeddings for land‑use/land‑cover mapping. By freezing the encoder and recomputing embeddings from a full year down to a single day, the authors benchmark linear probes and UNet heads on LUCAS, DynamicEarthNet, and PASTIS‑R datasets. Results show that embeddings are highly task‑dependent: for phenology‑driven classes (PASTIS‑R) they outperform from‑scratch models by ~46%, while for temporally stable classes (DynamicEarthNet, LUCAS) they match only with full supervision, yet remain more label‑efficient across all datasets.
By Julia Guerrero-Viu, Alex L\'opez-Cifuentes, Ignacio P\'erez-Villar, Fabio Pacifici
DistillPath-KS16 is a 22‑million‑parameter ViT‑S/16 pathology encoder distilled from larger teachers ranging from 86 M to 1.1 B parameters. By training only on the teachers’ final class and patch tokens across 6,000 public slides, it avoids costly pretraining heads and large tile corpora, yet surpasses the kaiko baseline on EVA, HEST, and PLISM benchmarks. The strongest variant, DistillPath-KS16‑Virchow2, achieves a mean EVA score of 0.795—just 0.015 points below the top model—while being 29× smaller and 25× faster.
arXiv:2609.36374v1 Announce Type: new
Abstract: Feed-forward 3D reconstruction models have achieved impressive performance by scaling model and dataset size, but their cost excludes most research gro...
By Brandon Leblanc, Charalambos Poullis
arXiv:2606. 27978v1 Announce Type: cross Abstract: Pixel-space continuous-token autoregressive (AR) generation directly models images as sequences of raw pixel patches, avoiding discrete tokenization or a separately pretrained tokenizer.
By Jiayi Xu, Di He, Guolin Ke
arXiv:2603. 19312v3 Announce Type: replace Abstract: Joint Embedding Predictive Architectures (JEPAs) offer a compelling framework for learning world models in compact latent spaces, yet existing methods remain fragile, relying on complex multi-term losses, exponential moving averages, pre-trained encoders, or auxiliary supervision to avoid representation collapse.
By Lucas Maes, Quentin Le Lidec, Damien Scieur, Yann LeCun, Randall Balestriero