SUFLECA is a weakly supervised framework that improves zero‑shot CAD‑to‑image alignment by scaling geometry‑grounded feature learning using Normalized Object Coordinates across up to 12 real and synthetic datasets. It introduces a geometrically consistent matching algorithm that reliably establishes CAD‑to‑image correspondences, enabling accurate, sub‑second alignment without iterative pose refinement. On the ScanNet25k benchmark, SUFLECA achieves 32.8%/42.6% category/instance accuracy, outperforming the strongest zero‑shot baseline by 9.7/12.5 percentage points and surpassing existing pose‑supervised methods for the first time.
By Saad Ejaz, Miguel Fernandez-Cortizas, Javier Civera, Holger Voos, Jose Luis Sanchez-Lopez
arXiv:2609.38578v1 Announce Type: cross
Abstract: Cross-structural motion retargeting aims to transfer motion between different skeletal topologies. Despite recent progress, existing state-of-the-art...
By Kia-J\"ung Yang, Fabian H. Sinz, Pawe{\l} A. Pierzchlewicz
TokenMatch is a transformer-based model that estimates 3D shape correspondences by adaptively tokenising meshes into curvature-guided patches. Trained only on the BeCoS partial-to-partial dataset, it generalises to full-shape matching without retraining, using self‑ and cross‑attention to learn patch‑ and point‑level relations. Evaluated on CP2P, PSMAL, BeCoS, FAUST, SCAPE, and SHREC'19, TokenMatch consistently outperforms existing methods in mean geodesic error and intersection‑over‑union while achieving sub‑second inference speeds.
By Adeela Islam, Zorah L\"ahner, Vittorio Murino, Vladislav Golyanik
arXiv:2608. 15710v1 Announce Type: cross Abstract: We address a fundamental gap in 3D-LLMs: existing models focus on single-object/scene description, struggling with detailed, inter-object comparison.
By Kohsuke Ide, Ryousuke Yamada, Yue Qiu, Xianzheng Ma, Yoshihiro Fukuhara, Hirokatsu Kataoka, Yutaka Satoh
arXiv:2602. 07429v2 Announce Type: replace-cross Abstract: Boundary representation (B-rep) is the industry standard for computer-aided design (CAD).
By Yuanxu Sun, Yuezhou Ma, Haixu Wu, Guanyang Zeng, Muye Chen, Jianmin Wang, Mingsheng Long
UniMate is a unified foundation model that generates articulated motion for any skeleton from a rigged 3D asset and a text prompt, eliminating the need for test‑time optimization or per‑skeleton retraining. It uses a topology‑aware diffusion transformer that incorporates skeletal topology through graph‑aware attention bias, spectral rotary position embedding, and a global topological conditioner. Trained on the newly curated UniML3D dataset of 13,006 diverse motion sequences, UniMate outperforms existing baselines in quality, generalization, and efficiency, and supports zero‑shot cross‑topology transfer, in‑betweening, expansion, and text‑guided editing.
By Linzhan Mou, Jiahui Lei, Zhiyang Dou, Chenyue Cai, Chaoyue Song, Adam Finkelstein, Szymon Rusinkiewicz
Vision Foundation Models (VFMs) have significantly advanced dense feature matching, yet severe in-plane rotation remains a critical challenge. Existing solutions face a fundamental dilemma: data-driven methods require inefficient parameter scaling to implicitly learn rotations, whereas strictly equivariant networks lack the semantic capacity of modern VFMs.
arXiv:2607. 00514v1 Announce Type: cross Abstract: Automatic understanding of dynamic 4D point clouds, the 3D-point sequences captured over time by depth sensors and LiDAR, is central to robotics and embodied perception.
By Trung Thanh Nguyen, Hai Nguyen-Truong, Tu Vo, Hoang M. Truong, Tuan-Anh Vu
PixelDense introduces a dual‑stream representation alignment for pixel diffusion, separating semantic and geometric teachers (DINOv2, SAM2, Depth Anything v2, Metric3D v2) into distinct projection spaces with an orthogonality penalty. The method improves dense‑prediction benchmarks, boosting PixelGen‑XXL’s GenEval score from 0.7927 to 0.8093, achieving significant gains in panoptic quality and depth accuracy, and accelerating training from random initialization. It also enhances SDEdit editing by preserving background structure and increasing PSNR.
By Lehan Yang, Daiqing Qi, Wenhao Zhang, Avery Li, Yiqing Yang, Yifan Li, Yu Kong, Haitian Zheng, Zhifei Zhang, Zhe Lin, Varun Jampani, Sheng Li
arXiv:2607. 17342v1 Announce Type: cross Abstract: Understanding physical human-robot and human-human interactions is a challenging yet emerging topic in 3D vision.
By Yuhang Wen, Mengyuan Liu, Zixuan Tang, Junsong Yuan, Sirui Li, Beichen Ding
arXiv:2609.12723v1 Announce Type: new
Abstract: Floorplans arise in many forms, from vector CAD drawings to raster renderings and sensor-derived density maps. This heterogeneity makes it difficult to...
By Xavier Anad\'on, R\'emi Pautrat, Rui Wang
arXiv:2409.11018v3 Announce Type: replace
Abstract: The LiDAR 3D object detector that balances accuracy and speed is crucial for achieving real-time perception in autonomous driving. However, many ex...
By Rui Yu, Runkai Zhao, Jiagen Li, Qingsong Zhao, HuaiCheng Yan, Meng Wang