arXiv Computer Vision By John Church, Vazghen Nikolian

Vision Foundation Models with Synthetic-Only Training for Monocular Spacecraft Pose Estimation

Read the original on arXiv Computer Vision →

The paper presents a new monocular spacecraft pose estimation model that achieves the lowest reported mean rotation errors on the SPEED+ lightbox and sunlamp test sets. By replacing smaller encoders with a large self‑supervised ViT foundation model (DINOv3) and scaling up to 840 M parameters, the authors improve accuracy from 300 M to 840 M parameters without saturation. The 840 M model also runs on a Jetson Orin NX 16 GB with 133.8 ms per crop and 32.0 W power draw, demonstrating embedded inference feasibility while training solely on synthetic data.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Machine Learning
Aug 28

Cross-simulator transfer with foundation model summaries: Towards robust SKA-era reionization inference

The paper demonstrates that a self‑supervised Vision Transformer (ViT) pretrained on a fast, low‑cost semi‑numerical simulator can produce data summaries that transfer across different simulators without retraining. In 21cm cosmology, the ViT—named SKATR—pretrained on 67,000 21cmFAST lightcones is applied unchanged to hydrodynamical Loreli II lightcones, enabling accurate inference of five astrophysical parameters with fewer radiative‑transfer simulations than a fully‑supervised baseline. SKATR remains accurate, informative, and calibrated even under realistic SKA antenna array noise, outperforming supervised models retrained on noisy data.

By Yannic Pietschke, Caroline Heneka, Ayodele Ore, Romain Meriot
arXiv Machine Learning
6d ago

Exploiting answer-invariant redundancies in satellite imagery for efficient VLM inference on edge

The paper introduces Rift, a two‑stage system that reduces the computational load of vision‑language models on satellites by pruning image tiles that do not affect the answer and then applying elastic prefill to limit token usage. By exploiting answer‑invariant token redundancy, Rift cuts energy consumption by 78 % and latency by 69 % compared to exhaustive tiled inference, while boosting accuracy from 45 % to 73 % on LLaVA‑1.5 7B running on a Jetson AGX Orin.

By Ishani Janveja, Davis Zhang, Seoyul Oh, Deepak Vasisht
Hugging Face Trending Papers
Jun 23

REDI-Match: Rotation-Equivariant Distillation for Efficient and Robust Dense Matching

Vision Foundation Models (VFMs) have significantly advanced dense feature matching, yet severe in-plane rotation remains a critical challenge. Existing solutions face a fundamental dilemma: data-driven methods require inefficient parameter scaling to implicitly learn rotations, whereas strictly equivariant networks lack the semantic capacity of modern VFMs.