arXiv AI

GAP-GDRNet: Geometry-Aware Monocular Visual Pose Sensing on a Single-Target Synthetic Spacecraft Dataset

arXiv:2607. 02360v1 Announce Type: cross Abstract: Monocular relative pose sensing is a central perception problem in non-cooperative rendezvous and on-orbit servicing.

arXiv Computer Vision
Sep 22

G6D: Geometric Learning-Free RGB-D 6D Pose Solver for Robotic Manipulation

G6D is a learning‑free, geometry‑driven RGB‑D 6D pose solver designed for robotic manipulation. It generates pose hypotheses via template‑based geometric matching and refines them using silhouette and depth consistency, requiring only an RGB‑D observation, an object mask, camera intrinsics, and a CAD model. The method offers adjustable accuracy‑computation trade‑offs, can run on CPU without GPUs, and has shown strong performance on LineMOD and BOP19 datasets, as well as in real‑world pick‑and‑place experiments.

By Yixuan Liang (Tsinghua University), William Chen (Sapient Intelligence), Yunan Wang (Tsinghua University), Jizhou Yan (Tsinghua University), Zhao Jin (Tsinghua University), Changling Liu (Sapient Intelligence), Chuxiong Hu (Tsinghua University)
arXiv Computer Vision
6d ago

GenCOPE: Syn2Real Generalized Category-Level Object Pose Estimation for Robotic Picking

GenCOPE introduces a synthetic-to-real (Syn2Real) approach for category-level object pose estimation (COPE) that eliminates the need for labor-intensive real-world data collection. By learning domain-invariant representations through 2D and 3D semantic consistency constraints and employing an end-to-end pose regression framework with 2D-3D cross consistency, the model achieves robust generalization across synthetic and real domains. The architecture relies solely on global features, resulting in a lightweight and efficient design validated on REAL275, Wild6D, and real-world robotic manipulation scenes.

By Jian Liu, Wei Sun, Zhenqi Dai, Hui Yang, Jian Xiao, Nicu Sebe, Na Zhao
arXiv Computer Vision
3d ago

RYOPO: Bringing End-to-End Category-Level Object Pose Estimation into Real Time

RYOPO is an end‑to‑end query‑based RGB‑D set predictor that jointly detects, segments, and estimates 9‑DoF poses of unseen instances within known categories without relying on external instance segmentation or CAD priors. It uses shared image and scene encoding, a query‑conditioned geometry pathway, and object‑centric refinement with pose‑conditioned cross‑attention to achieve accurate pose estimation. On benchmark datasets such as NOCS, REAL275, and HouseCat6D, RYOPO outperforms published methods and runs in real time at 31.8 FPS on an RTX A6000.

By Hakjin Lee, Junghoon Seo, Jaehoon Sim
arXiv Computer Vision
Sep 7

CrossDepth: Geometry-Constrained Attention for Generalizable Multi-View Surround Depth Estimation

CrossDepth introduces geometry-constrained attention for multi-view surround depth estimation, addressing cross-image inconsistencies caused by varying camera intrinsics and limited receptive fields. The method conditions features on per-pixel camera-aware ray embeddings and extends pixel context via cross-image attention limited to geometrically plausible regions. Trained self-supervised with photometric consistency, it achieves better depth accuracy and consistency on DDAD and nuScenes compared to existing self-supervised approaches.

By Samer Abualhanud, Max Mehltretter
arXiv Computer Vision
Sep 18

G2G: Exploiting Intra-Group Geometry for Inter-Group Pose Estimation

The paper introduces G2G, a method that leverages known intra-group geometry to estimate the relative 6-DoF pose between two image groups, a key problem in cross-sequence relocalization and multi-camera rig odometry. G2G keeps a frozen foundation model and adds three lightweight trainable modules—a perceiver resampler, a cross-group bridge with merged self-attention, and a multi-frame pose head—totaling about 32 M parameters, less than 6% of the full model. Evaluated on four diverse datasets covering indoor/outdoor simulation, real-world cross-season capture, and zero-shot sim-to-real transfer, G2G achieves state‑of‑the‑art accuracy on both pose estimation tasks while only requiring supervision from relative poses.

By Yufei Wei, Shuhao Ye, Chenxiao Hu, Yiyuan Pan, Dongyu Feng, Rong Xiong, Yue Wang, Yanmei Jiao
arXiv Computer Vision
3d ago

SurGe: Improved Surface Geometry in Point Maps

arXiv:2605.31577v2 Announce Type: replace Abstract: Recent feedforward 3D reconstruction methods predict point maps and estimate global 3D geometry remarkably well. However, their predictions still e...

By Karim Knaebel, Gonzalo Martin Garcia, Christian Schmidt, Ilya Fradlin, Lucas Nunes, Daan de Geus, Bastian Leibe
arXiv Computer Vision
Sep 3

Geometric Distillation from Rectified Stereo: Leveraging Epipolar Cues for Monocular Depth

The paper introduces Epipolar Distillation (EpiDistill), a method that transfers scale‑aware geometric priors from multi‑view models to monocular depth foundation models using Rectified Stereo Tokens. By preserving epipolar attention patterns, the single‑view model maintains geometric consistency without needing multi‑view inputs during inference. Experiments show significant improvements in zero‑shot metric depth estimation on challenging datasets such as ETH3D and DIODE, and the approach consistently boosts performance of state‑of‑the‑art ViT‑based models like UniDepthV2 and DepthPro.

By Jung-Hee Kim, Xiaoming Liu