arXiv AI By Zongwu Xie, Yonglong Zhang, Yifan Yang, Yang Liu, Guanghu Xie

GAP-GDRNet: Geometry-aware monocular 6D pose estimation for spacecraft using synthetic geometric supervision

Read the original on arXiv AI →

arXiv:2607. 02360v3 Announce Type: replace-cross Abstract: Monocular spacecraft 6D pose estimation remains difficult under weak texture, thin structures, illumination variation, and occlusion.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Sep 22

G6D: Geometric Learning-Free RGB-D 6D Pose Solver for Robotic Manipulation

G6D is a learning‑free, geometry‑driven RGB‑D 6D pose solver designed for robotic manipulation. It generates pose hypotheses via template‑based geometric matching and refines them using silhouette and depth consistency, requiring only an RGB‑D observation, an object mask, camera intrinsics, and a CAD model. The method offers adjustable accuracy‑computation trade‑offs, can run on CPU without GPUs, and has shown strong performance on LineMOD and BOP19 datasets, as well as in real‑world pick‑and‑place experiments.

By Yixuan Liang (Tsinghua University), William Chen (Sapient Intelligence), Yunan Wang (Tsinghua University), Jizhou Yan (Tsinghua University), Zhao Jin (Tsinghua University), Changling Liu (Sapient Intelligence), Chuxiong Hu (Tsinghua University)
arXiv Computer Vision
3d ago

RYOPO: Bringing End-to-End Category-Level Object Pose Estimation into Real Time

RYOPO is an end‑to‑end query‑based RGB‑D set predictor that jointly detects, segments, and estimates 9‑DoF poses of unseen instances within known categories without relying on external instance segmentation or CAD priors. It uses shared image and scene encoding, a query‑conditioned geometry pathway, and object‑centric refinement with pose‑conditioned cross‑attention to achieve accurate pose estimation. On benchmark datasets such as NOCS, REAL275, and HouseCat6D, RYOPO outperforms published methods and runs in real time at 31.8 FPS on an RTX A6000.

By Hakjin Lee, Junghoon Seo, Jaehoon Sim
arXiv Computer Vision
6d ago

GenCOPE: Syn2Real Generalized Category-Level Object Pose Estimation for Robotic Picking

GenCOPE introduces a synthetic-to-real (Syn2Real) approach for category-level object pose estimation (COPE) that eliminates the need for labor-intensive real-world data collection. By learning domain-invariant representations through 2D and 3D semantic consistency constraints and employing an end-to-end pose regression framework with 2D-3D cross consistency, the model achieves robust generalization across synthetic and real domains. The architecture relies solely on global features, resulting in a lightweight and efficient design validated on REAL275, Wild6D, and real-world robotic manipulation scenes.

By Jian Liu, Wei Sun, Zhenqi Dai, Hui Yang, Jian Xiao, Nicu Sebe, Na Zhao