The paper introduces a CAD‑free 3D shape prior that enhances object recognition by reconstructing each object with 3D Gaussian Splatting (3DGS) from short RGB‑D scans and fusing the resulting shape prototype with frozen DINOv2 image features. Experiments on T‑LESS and HOPE datasets show that geometry alone can match or exceed CAD‑based recognition, and that the combined approach improves performance, especially on shape‑distinctive or partially occluded objects. The study demonstrates that the benefit comes from the geometric information rather than rendered pixels, and that the prior is complementary to frozen vision features.
By Chenxi Tao, Seung-Kyum Choi
arXiv:2609.13263v1 Announce Type: new
Abstract: Projective shape analysis provides a geometric framework for studying landmark configurations in digital images acquired by pinhole cameras. In the cla...
By Musab Alamoudi, Robert L. Paige, Vic Patrangenaru
arXiv:2609.26795v1 Announce Type: cross
Abstract: 3D Gaussian Splatting (3DGS) can reconstruct a captured scene photorealistically, but the resulting representation does not by itself support physica...
By Runyi Yang, Deheng Zhang, Xiaoye Wang, Kanzhi Wu, Lei Sun, Ajad Chhatkuli, Kunyu Peng, Luc Van Gool, Danda Pani Paudel
The paper introduces SE(3) neural potential fields that learn collision‑free 6‑DoF trajectory planning directly from posed RGB images, eliminating the need for explicit 3D reconstruction. By supervising the field with a navigation function based on geodesic distances to the grasp, the method avoids the classic pitfalls of artificial potential fields, achieving near‑goal convergence within 3 cm from any start and producing collision‑free paths on a UR10 robot. Experiments on two tabletop scenes show significant improvements in clearance, reduced arm‑link contacts, and a 90 % grasp success rate, while planning time drops from over a minute to about 2 seconds compared to RRT* on a reconstructed scene.
By Jeffrey Eiyike, Masoud Ataei, Elvis Gyaase, Vikas Dhiman
CADSplat is a framework that reconstructs photorealistic, geometrically accurate digital twins from fewer than 15 wide‑baseline images by regularizing 3D Gaussian Splatting with an explicit CAD shape prior. It matches segmented object silhouettes to a CAD library to retrieve a suitable model and camera poses, then anchors Gaussian primitives to the model’s surface and jointly optimizes splat parameters, registration, and a non‑rigid deformation field. Experiments on two real‑world datasets show CADSplat outperforms baselines, especially in sparse and self‑occluded scenarios, and its gains mainly stem from constraining splats to a surface rather than the CAD shape itself.
By Kristof Overdulve, Lode Jorissen, Nick Michiels
Reaching a 6-DoF grasp pose in clutter requires a collision-free trajectory, conventionally obtained by reconstructing the scene in 3D and planning inside that reconstruction, at the cost of its accur...
arXiv:2608.31002v1 Announce Type: cross
Abstract: Robotic perception from a single viewpoint is often limited by self-occlusion and incomplete surface visibility. This paper presents DARP(Dual-Arm Ro...
By Manish Kansana, Mohammed Yusuf Mujawar, Sudip Mittal, Shahram Rahimi, Noorbakhsh Amiri Golilarz
arXiv:2609.23182v1 Announce Type: new
Abstract: Feed-forward 3D Gaussian Splatting now reconstructs renderable scenes from unposed, uncalibrated images. Yet, most models supervise only photometric co...
By Si-Yu Lu, Yung-Yao Chen, Yi Jan Chen, Shang-Lin Li, Ching-Chan Liao, Wen-Huang Cheng
AnyBox is a zero‑shot framework that estimates the full 9DoF pose (6D pose plus 3D dimensions) of boxes from a single RGB‑D image, leveraging the geometric regularity of boxes. It alternates between pose and scale estimation, using a binary search guided by the discrepancy between a reprojected template and the observed mask, and employs a depth‑consistency filter and an early‑stopping rule to prune implausible hypotheses. On public benchmarks and a warehouse dataset, AnyBox improves detection AP by up to 36 points and boosts robotic box‑shelving success by 28%.
By Yintao Ma, Sajjad Pakdamansavoji, Charles Eret, Rui Heng Yang, Xuan Zhao, Yingxue Zhang, Tongtong Cao, Amir Rasouli
Matisse is a training‑free framework that combines active 3D reconstruction with keyframe selection by using evidence from a pretrained generative 3D model. It estimates evidential uncertainty via cross‑attention on 3D latent tokens and derives an evidential information gain to guide view acquisition and keyframe selection, reducing redundant observations and supporting multi‑object scenes with occlusion‑aware aggregation. On GSO30, YCB‑V, and Replica, Matisse improves Chamfer distance by 12.7%, 3.8%, and 9.2% respectively, and speeds up end‑to‑end reconstruction by 1.5× compared to the best baseline.
By Xihang Yu, Kaichen Zhou, Lorenzo Shaikewitz, Cl\'ement Jambon, Xiao Zhan, Rajat Talak, Luca Carlone
arXiv:2608.30521v1 Announce Type: new
Abstract: Conventional algebraic triangulation solves 3D human pose estimation (HPE) from multi-view 2D keypoints. The typical approach, decoding 2D keypoints fr...
By Ziliang Xiong, Henglin Shi, Per-Erik Forssen
3D Gaussian Splatting (3DGS) can reconstruct a captured scene photorealistically, but the resulting representation does not by itself support physical interaction. Robot simulation instead requires ob...