RYOPO is an end‑to‑end query‑based RGB‑D set predictor that jointly detects, segments, and estimates 9‑DoF poses of unseen instances within known categories without relying on external instance segmentation or CAD priors. It uses shared image and scene encoding, a query‑conditioned geometry pathway, and object‑centric refinement with pose‑conditioned cross‑attention to achieve accurate pose estimation. On benchmark datasets such as NOCS, REAL275, and HouseCat6D, RYOPO outperforms published methods and runs in real time at 31.8 FPS on an RTX A6000.
By Hakjin Lee, Junghoon Seo, Jaehoon Sim
MessyKitchens introduces a new dataset of cluttered real-world kitchen scenes with detailed 3D object shapes, poses, and accurate contact information. The authors extend the SAM 3D single-object reconstruction method with a Multi-Object Decoder (MOD) to jointly reconstruct entire scenes, achieving better registration accuracy and reduced inter-object penetration compared to prior work. The dataset, benchmark, code, and pretrained models will be publicly released on the project website.
By Junaid Ahmed Ansari, Ran Ding, Fabio Pizzati, Ivan Laptev
arXiv:2605.13018v2 Announce Type: replace
Abstract: Object-centric scene understanding is a fundamental challenge in computer vision. Existing approaches often rely on multi-stage pipelines that firs...
By Yi Du, Yang You, Xiang Wan, Leonidas Guibas
Category-level object pose estimation seeks to recover a similarity transform $(R,t,s)$ for unseen instances without instance-specific CAD models. Most competitive methods are correspondence-based: pr...
PriorPose introduces a reference-guided correspondence framework for category-level object pose estimation that jointly performs canonicalization and alignment in a shared feature space. By embedding partial observations and a category prior as token sets in a seeded transformer, the network predicts per-point NOCS fields and a canonical deformation, while a deep pose head regresses the similarity transform. A two-part shape consistency objective couples correspondence, deformation, and pose, reducing reliance on memorized canonical orientations and avoiding error cascades, leading to state-of-the-art results on standard and larger-category benchmarks, especially under strict pose thresholds.
By Yihan Chen, Huan Ren, Wenfei Yang, Hang Du, Tianzhu Zhang, Feng Wu
SUFLECA is a weakly supervised framework that improves zero‑shot CAD‑to‑image alignment by scaling geometry‑grounded feature learning using Normalized Object Coordinates across up to 12 real and synthetic datasets. It introduces a geometrically consistent matching algorithm that reliably establishes CAD‑to‑image correspondences, enabling accurate, sub‑second alignment without iterative pose refinement. On the ScanNet25k benchmark, SUFLECA achieves 32.8%/42.6% category/instance accuracy, outperforming the strongest zero‑shot baseline by 9.7/12.5 percentage points and surpassing existing pose‑supervised methods for the first time.
By Saad Ejaz, Miguel Fernandez-Cortizas, Javier Civera, Holger Voos, Jose Luis Sanchez-Lopez
arXiv:2607. 04930v1 Announce Type: cross Abstract: In the pursuit of robust and generalizable category-level object pose estimation, most existing methods adopt parametric formulations that learn effective representations from data, yet they primarily encode category-level patterns into fixed shape priors or static parameter weights, which limits their scalability to highly diverse instances.
By Xiao Lin, Minghao Zhu, Yun Peng, Liuyi Wang, Qiyi Wang, Chengju Liu, Qijun Chen
GenCOPE introduces a synthetic-to-real (Syn2Real) approach for category-level object pose estimation (COPE) that eliminates the need for labor-intensive real-world data collection. By learning domain-invariant representations through 2D and 3D semantic consistency constraints and employing an end-to-end pose regression framework with 2D-3D cross consistency, the model achieves robust generalization across synthetic and real domains. The architecture relies solely on global features, resulting in a lightweight and efficient design validated on REAL275, Wild6D, and real-world robotic manipulation scenes.
By Jian Liu, Wei Sun, Zhenqi Dai, Hui Yang, Jian Xiao, Nicu Sebe, Na Zhao
The paper presents a generative framework that estimates category-level 6D pose and 3D size of objects from a single RGB image, using score-based diffusion models to produce a multi-hypothesis pose distribution. It replaces costly likelihood pruning with a Mean Shift approach to isolate the mode as the final pose estimate, achieving state-of-the-art results on the REAL275 benchmark. The method also decouples detection from pose estimation, enabling robust zero-shot generalisation on the Wild6D dataset and extending naturally to video sequences by propagating the pose distribution over time.
By Adam Bethell, Ravi Garg, Ian Reid
arXiv:2609.39116v1 Announce Type: new
Abstract: Prior-free 6D object pose tracking seeks to recover the trajectory of an unseen object from a single RGB video without object-specific CAD models, pose...
By Shiyang Liu, Weiquan Lin, Luping Xiao, Jiadong Tang, Yi Yang, Yu Gao, Xingyu Chen
arXiv:2511. 16624v2 Announce Type: replace-cross Abstract: We present SAM 3D, a generative model for visually grounded 3D object reconstruction, predicting geometry, texture, and layout from a single image.
By SAM 3D Team, Xingyu Chen, Fu-Jen Chu, Pierre Gleize, Kevin J Liang, Alexander Sax, Hao Tang, Weiyao Wang, Michelle Guo, Thibaut Hardin, Xiang Li, Aohan Lin, Jiawei Liu, Ziqi Ma, Anushka Sagar, Bowen Song, Xiaodong Wang, Jianing Yang, Bowen Zhang, Piotr Doll\'ar, Georgia Gkioxari, Matt Feiszli, Jitendra Malik
In the pursuit of robust and generalizable category-level object pose estimation, most existing methods adopt parametric formulations that learn effective representations from data, yet they primarily encode category-level patterns into fixed shape priors or static parameter weights, which limits their scalability to highly diverse instances. In this paper, we rethink category-level pose estimation from a memory-centric perspective and present MemPose, a memory-augmented framework that explicitly incorporates category-level geometric memory into the pose estimation pipeline.