arXiv:2607. 02360v3 Announce Type: replace-cross Abstract: Monocular spacecraft 6D pose estimation remains difficult under weak texture, thin structures, illumination variation, and occlusion.
By Zongwu Xie, Yonglong Zhang, Yifan Yang, Yang Liu, Guanghu Xie
arXiv:2609.22868v1 Announce Type: new
Abstract: End-to-end driving requires planning-relevant bird's-eye-view (BEV) representations, but existing pretraining approaches often rely on task annotations...
By Jaeha Song, Soonmin Hwang
arXiv:2609.21597v1 Announce Type: new
Abstract: Monocular 6-DoF pose estimation of non-cooperative targets is important for on-orbit servicing and debris removal. A single-image estimator can confuse...
By Andr\'e Lopo, Atabak Dehban, Rodrigo Ventura
The paper demonstrates that a self‑supervised Vision Transformer (ViT) pretrained on a fast, low‑cost semi‑numerical simulator can produce data summaries that transfer across different simulators without retraining. In 21cm cosmology, the ViT—named SKATR—pretrained on 67,000 21cmFAST lightcones is applied unchanged to hydrodynamical Loreli II lightcones, enabling accurate inference of five astrophysical parameters with fewer radiative‑transfer simulations than a fully‑supervised baseline. SKATR remains accurate, informative, and calibrated even under realistic SKA antenna array noise, outperforming supervised models retrained on noisy data.
By Yannic Pietschke, Caroline Heneka, Ayodele Ore, Romain Meriot
The paper introduces Rift, a two‑stage system that reduces the computational load of vision‑language models on satellites by pruning image tiles that do not affect the answer and then applying elastic prefill to limit token usage. By exploiting answer‑invariant token redundancy, Rift cuts energy consumption by 78 % and latency by 69 % compared to exhaustive tiled inference, while boosting accuracy from 45 % to 73 % on LLaVA‑1.5 7B running on a Jetson AGX Orin.
By Ishani Janveja, Davis Zhang, Seoyul Oh, Deepak Vasisht
Vision Foundation Models (VFMs) have significantly advanced dense feature matching, yet severe in-plane rotation remains a critical challenge. Existing solutions face a fundamental dilemma: data-driven methods require inefficient parameter scaling to implicitly learn rotations, whereas strictly equivariant networks lack the semantic capacity of modern VFMs.
arXiv:2609.38755v1 Announce Type: new
Abstract: A wide range of approaches have been developed for camera pose estimation, including correspondence-based methods, end-to-end pose regression, and rece...
By Zhining Gu, Shangjie Du, Weimin Qiu, Carl Olsson, Ping Liu, Meng Tang
The study investigates how to automatically predict the build orientation for selective laser melting (SLM) of dental parts using supervised machine learning. Researchers trained two different neural network backbones—ResNet‑50 on multi‑view images and PointNeXt‑S on point clouds—on about 2,400 patient‑specific parts, evaluating 13 different ways to represent the up‑axis (six classical SO(3) parameterizations and seven unit‑sphere representations). They found that applying test‑time augmentation (TTA) over 21 known rotations consistently reduced angular error, with the octahedral map achieving the lowest mean error (10.6°) on ResNet‑50, while direct S² representations performed best overall but may be influenced by label noise.
By Felix Schmalzel, Reimar Waitz, Moritz Kronberger, Thorsten Sch\"oler
The paper introduces CalfVO, a monocular visual odometry system that operates without camera intrinsics, test‑time optimization, bundle adjustment, or loop closure. Using a transformer, it predicts relative poses with separate rotation and translation confidences over overlapping image windows, then aggregates these predictions via a confidence‑weighted module to produce a single trajectory. CalfVO achieves the highest accuracy among calibration‑free methods across five benchmarks and runs at 53 FPS, outperforming all baselines.
By Vladimir Yugay, Duy-Kien Nguyen, Theo Gevers, Cees G. M. Snoek, Martin R. Oswald
The paper introduces OVIE, a monocular novel-view synthesis method that eliminates the need for multi‑view training data. By using a frozen depth estimator to generate pseudo‑target views from single images and applying masked and adversarial losses, OVIE is trained on 30 million uncurated images. It achieves state‑of‑the‑art performance on RealEstate10K and DL3DV, produces highly consistent multi‑view trajectories, and runs at 116 FPS—over 600× faster than the fastest baseline.
By Adrien Ramanana Rahary, Nicolas Dufour, Patrick Perez, David Picard
arXiv:2504.15776v2 Announce Type: replace
Abstract: Public autonomous driving datasets underpin the training and benchmarking of perception, mapping, and localization algorithms, yet residual inaccur...
By Quentin Herau, Nathan Piasco, Moussab Bennehar, Luis Rold\~ao, Dzmitry Tsishkou, Bingbing Liu, Cyrille Migniot, Pascal Vasseur, C\'edric Demonceaux
arXiv:2607. 17099v1 Announce Type: cross Abstract: Recent geometric foundation models (e.
By Feng Xue, Wu Chen, Mingshuai Zhao, Guofeng Zhong, Anlong Ming, Haozhe Wang, Dianqiao Lei, Zhaowen Lin, Haiyang Zhang, Nicu Sebe