In the pursuit of robust and generalizable category-level object pose estimation, most existing methods adopt parametric formulations that learn effective representations from data, yet they primarily encode category-level patterns into fixed shape priors or static parameter weights, which limits their scalability to highly diverse instances. In this paper, we rethink category-level pose estimation from a memory-centric perspective and present MemPose, a memory-augmented framework that explicitly incorporates category-level geometric memory into the pose estimation pipeline.
arXiv:2610.03013v1 Announce Type: new
Abstract: Category-level object pose estimation predicts the rotation, translation, and metric size of unseen instances within known categories. Many accurate RG...
By Hakjin Lee, Junghoon Seo, Jaehoon Sim
PriorPose introduces a reference-guided correspondence framework for category-level object pose estimation that jointly performs canonicalization and alignment in a shared feature space. By embedding partial observations and a category prior as token sets in a seeded transformer, the network predicts per-point NOCS fields and a canonical deformation, while a deep pose head regresses the similarity transform. A two-part shape consistency objective couples correspondence, deformation, and pose, reducing reliance on memorized canonical orientations and avoiding error cascades, leading to state-of-the-art results on standard and larger-category benchmarks, especially under strict pose thresholds.
By Yihan Chen, Huan Ren, Wenfei Yang, Hang Du, Tianzhu Zhang, Feng Wu
Category-level object pose estimation seeks to recover a similarity transform $(R,t,s)$ for unseen instances without instance-specific CAD models. Most competitive methods are correspondence-based: pr...
GenCOPE introduces a synthetic-to-real (Syn2Real) approach for category-level object pose estimation (COPE) that eliminates the need for labor-intensive real-world data collection. By learning domain-invariant representations through 2D and 3D semantic consistency constraints and employing an end-to-end pose regression framework with 2D-3D cross consistency, the model achieves robust generalization across synthetic and real domains. The architecture relies solely on global features, resulting in a lightweight and efficient design validated on REAL275, Wild6D, and real-world robotic manipulation scenes.
By Jian Liu, Wei Sun, Zhenqi Dai, Hui Yang, Jian Xiao, Nicu Sebe, Na Zhao
Moving6DPoSe is a multimodal database for monocular 6D pose estimation and segmentation of moving objects, comprising two subsets: real-world recordings (Moving6DPoSe‑R) and synthetic sequences (Moving6DPoSe‑S). It includes 16 scanned objects, 1,702 real and synthetic rosbags, and annotations for semantic segmentation, object detection, and monocular 6D pose estimation. Baseline results show that event-based representations outperform conventional RGB images for moving‑object segmentation, while monocular orientation estimation remains challenging.
By Ignacio Bugueno-Cordova, Javier Ruiz-del-Solar, Rodrigo Verschae