arXiv:2607. 02360v3 Announce Type: replace-cross Abstract: Monocular spacecraft 6D pose estimation remains difficult under weak texture, thin structures, illumination variation, and occlusion.
By Zongwu Xie, Yonglong Zhang, Yifan Yang, Yang Liu, Guanghu Xie
arXiv:2609.22868v1 Announce Type: new
Abstract: End-to-end driving requires planning-relevant bird's-eye-view (BEV) representations, but existing pretraining approaches often rely on task annotations...
By Jaeha Song, Soonmin Hwang
arXiv:2609.21597v1 Announce Type: new
Abstract: Monocular 6-DoF pose estimation of non-cooperative targets is important for on-orbit servicing and debris removal. A single-image estimator can confuse...
By Andr\'e Lopo, Atabak Dehban, Rodrigo Ventura
The paper demonstrates that a self‑supervised Vision Transformer (ViT) pretrained on a fast, low‑cost semi‑numerical simulator can produce data summaries that transfer across different simulators without retraining. In 21cm cosmology, the ViT—named SKATR—pretrained on 67,000 21cmFAST lightcones is applied unchanged to hydrodynamical Loreli II lightcones, enabling accurate inference of five astrophysical parameters with fewer radiative‑transfer simulations than a fully‑supervised baseline. SKATR remains accurate, informative, and calibrated even under realistic SKA antenna array noise, outperforming supervised models retrained on noisy data.
By Yannic Pietschke, Caroline Heneka, Ayodele Ore, Romain Meriot
The paper introduces Rift, a two‑stage system that reduces the computational load of vision‑language models on satellites by pruning image tiles that do not affect the answer and then applying elastic prefill to limit token usage. By exploiting answer‑invariant token redundancy, Rift cuts energy consumption by 78 % and latency by 69 % compared to exhaustive tiled inference, while boosting accuracy from 45 % to 73 % on LLaVA‑1.5 7B running on a Jetson AGX Orin.
By Ishani Janveja, Davis Zhang, Seoyul Oh, Deepak Vasisht
Vision Foundation Models (VFMs) have significantly advanced dense feature matching, yet severe in-plane rotation remains a critical challenge. Existing solutions face a fundamental dilemma: data-driven methods require inefficient parameter scaling to implicitly learn rotations, whereas strictly equivariant networks lack the semantic capacity of modern VFMs.