arXiv Computer Vision
3d ago

RYOPO: Bringing End-to-End Category-Level Object Pose Estimation into Real Time

RYOPO is an end‑to‑end query‑based RGB‑D set predictor that jointly detects, segments, and estimates 9‑DoF poses of unseen instances within known categories without relying on external instance segmentation or CAD priors. It uses shared image and scene encoding, a query‑conditioned geometry pathway, and object‑centric refinement with pose‑conditioned cross‑attention to achieve accurate pose estimation. On benchmark datasets such as NOCS, REAL275, and HouseCat6D, RYOPO outperforms published methods and runs in real time at 31.8 FPS on an RTX A6000.

By Hakjin Lee, Junghoon Seo, Jaehoon Sim
arXiv AI
Sep 24

MessyKitchens: Contact-rich object-level 3D scene reconstruction

MessyKitchens introduces a new dataset of cluttered real-world kitchen scenes with detailed 3D object shapes, poses, and accurate contact information. The authors extend the SAM 3D single-object reconstruction method with a Multi-Object Decoder (MOD) to jointly reconstruct entire scenes, achieving better registration accuracy and reduced inter-object penetration compared to prior work. The dataset, benchmark, code, and pretrained models will be publicly released on the project website.

By Junaid Ahmed Ansari, Ran Ding, Fabio Pizzati, Ivan Laptev
arXiv Computer Vision
Sep 16

PriorPose: Reference-Guided Joint Deformation and Alignment for Category-Level Object Pose Estimation

PriorPose introduces a reference-guided correspondence framework for category-level object pose estimation that jointly performs canonicalization and alignment in a shared feature space. By embedding partial observations and a category prior as token sets in a seeded transformer, the network predicts per-point NOCS fields and a canonical deformation, while a deep pose head regresses the similarity transform. A two-part shape consistency objective couples correspondence, deformation, and pose, reducing reliance on memorized canonical orientations and avoiding error cascades, leading to state-of-the-art results on standard and larger-category benchmarks, especially under strict pose thresholds.

By Yihan Chen, Huan Ren, Wenfei Yang, Hang Du, Tianzhu Zhang, Feng Wu
arXiv Computer Vision
Aug 31

SUFLECA: Scaling Up Feature Learning for CAD-to-image Alignment

SUFLECA is a weakly supervised framework that improves zero‑shot CAD‑to‑image alignment by scaling geometry‑grounded feature learning using Normalized Object Coordinates across up to 12 real and synthetic datasets. It introduces a geometrically consistent matching algorithm that reliably establishes CAD‑to‑image correspondences, enabling accurate, sub‑second alignment without iterative pose refinement. On the ScanNet25k benchmark, SUFLECA achieves 32.8%/42.6% category/instance accuracy, outperforming the strongest zero‑shot baseline by 9.7/12.5 percentage points and surpassing existing pose‑supervised methods for the first time.

By Saad Ejaz, Miguel Fernandez-Cortizas, Javier Civera, Holger Voos, Jose Luis Sanchez-Lopez