arXiv Computer Vision

KASALv2: Fully Automatic 3D Rotational Symmetry Classification and Axis Localization

arXiv Computer Vision
1d ago

WAPR: A Foundation Model for Wide-Angle Refinement in Unseen Object Pose Estimation

WAPR is a zero‑shot wide‑angle pose refinement model that can correct candidate 6D poses with rotational errors up to 90°, achieving fast inference (≤1 s per frame) and high throughput (≈25 detections per second). It leverages rotational symmetry priors to canonicalize pose targets and introduces the SA6D dataset, which augments 944 GSO scans into ~50 k object instances and ~2 M RGB‑D images. Experiments on seven BOP core datasets demonstrate that WAPR sets new state‑of‑the‑art performance for unseen‑object pose estimation in both fast and unconstrained settings.

By Yulin Wang, Mengting Hu, Hongli Li, Jianghao Zhou, Chen Luo
arXiv Computer Vision
Aug 28

A Geometry-Driven, Framework-Agnostic Optimization for Object Pose Estimation

The paper proposes a data‑centric optimization for object pose estimation that uses a physically grounded rotation representation based on principal axes alignment. By aligning an object's coordinate system with its inertial principal axes, the method achieves inherent stability, symmetry‑aware canonicalization, and framework agnosticism, allowing it to be applied at the dataset level without modifying existing networks. Experiments on category‑level and instance‑level models show consistent accuracy improvements while preserving baseline network integrity.

By Wei Chen, Tao Zhen, Zhongchen Shi, Jing Zhang, Liang Xie, Erwei Yin
arXiv Computer Vision
Sep 15

MoCapAnything V2: End-to-End Motion Capture for Arbitrary Skeletons

arXiv:2604.28130v4 Announce Type: replace Abstract: Recent methods for arbitrary-skeleton motion capture from monocular video follow a factorized pipeline, where a Video-to-Pose network predicts join...

By Kehong Gong, Zhengyu Wen, Dao Thien Phong, Mingxi Xu, Weixia He, Qi Wang, Ning Zhang, Zhengyu Li, Guanli Hou, Dongze Lian, Xiaoyu He, Mingyuan Zhang, Hanwang Zhang
arXiv Machine Learning
1d ago

How Do Transformers Learn to Represent Symmetries?

The paper investigates how a vanilla Transformer learns symmetries from finite data augmentation on point cloud datasets. It finds an ordering of learnability: non-angle-preserving symmetries are easiest, followed by angle-preserving symmetries, and finally base angle-preserving subgroups such as translation, rotation, and scale. The study also examines the Transformer's extrapolation behavior, performs a structural analysis of trained models to uncover interpretable mechanisms for invariance, and extends these findings to equivariant functions, suggesting that the identified mechanisms can serve as building blocks for learned equivariance.

By Eduardo Santos-Escriche, Valerie Engelmayer, Ya-Wei Eileen Lin, Stefanie Jegelka
arXiv Computer Vision
Sep 14

Spectral Consistency-Guided Multiview Point Cloud Registration for Low-Overlap Scenes

The paper introduces GMPCR, a non‑learning spectral consistency‑guided framework for multiview point cloud registration in low‑overlap scenes. GMPCR refines initial correspondences into a second‑order compatibility structure, uses spectral analysis to filter unreliable matches and select informative scan pairs, and then applies maximal‑clique hypothesis generation for robust relative transformations. The resulting sparse pose graph is further refined with an adaptive history‑aware synchronization scheme, and a recovery mechanism allows previously down‑weighted edges to regain confidence, achieving high registration recalls on benchmark datasets while reducing computational cost.

By Tianyu Li, Yanghong Lin, Shudong Zhou, Kui Yang, Jingru Zhang, Li Fang, Wei Yao
arXiv Computer Vision
Sep 16

PriorPose: Reference-Guided Joint Deformation and Alignment for Category-Level Object Pose Estimation

PriorPose introduces a reference-guided correspondence framework for category-level object pose estimation that jointly performs canonicalization and alignment in a shared feature space. By embedding partial observations and a category prior as token sets in a seeded transformer, the network predicts per-point NOCS fields and a canonical deformation, while a deep pose head regresses the similarity transform. A two-part shape consistency objective couples correspondence, deformation, and pose, reducing reliance on memorized canonical orientations and avoiding error cascades, leading to state-of-the-art results on standard and larger-category benchmarks, especially under strict pose thresholds.

By Yihan Chen, Huan Ren, Wenfei Yang, Hang Du, Tianzhu Zhang, Feng Wu