Google AI Blog By Google AI

MELON: Reconstructing 3D objects from images with unknown poses

Read the original on Google AI Blog →

Posted by Mark Matthews, Senior Software Engineer, and Dmitry Lagun, Research Scientist, Google Research A person's prior experience and understanding of the world generally enables them to easily infer what an object looks like in whole, even if only looking at a few 2D pictures of it. Yet the capacity for a computer to reconstruct the shape of an object in 3D given only a few images has remained a difficult algorithmic problem for years.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Google AI Blog.

arXiv AI
Sep 7

Where Appearance Fails, Geometry Recognizes: A CAD-Free 3D Shape Prior That Complements Vision Foundation Models

The paper introduces a CAD‑free 3D shape prior that enhances object recognition by reconstructing each object with 3D Gaussian Splatting (3DGS) from short RGB‑D scans and fusing the resulting shape prototype with frozen DINOv2 image features. Experiments on T‑LESS and HOPE datasets show that geometry alone can match or exceed CAD‑based recognition, and that the combined approach improves performance, especially on shape‑distinctive or partially occluded objects. The study demonstrates that the benefit comes from the geometric information rather than rendered pixels, and that the prior is complementary to frozen vision features.

By Chenxi Tao, Seung-Kyum Choi
arXiv AI
Sep 23

SE(3) Neural Potential Fields for 6-DoF Trajectory Planning Directly from Images Without Explicit 3D Reconstruction

The paper introduces SE(3) neural potential fields that learn collision‑free 6‑DoF trajectory planning directly from posed RGB images, eliminating the need for explicit 3D reconstruction. By supervising the field with a navigation function based on geodesic distances to the grasp, the method avoids the classic pitfalls of artificial potential fields, achieving near‑goal convergence within 3 cm from any start and producing collision‑free paths on a UR10 robot. Experiments on two tabletop scenes show significant improvements in clearance, reduced arm‑link contacts, and a 90 % grasp success rate, while planning time drops from over a minute to about 2 seconds compared to RRT* on a reconstructed scene.

By Jeffrey Eiyike, Masoud Ataei, Elvis Gyaase, Vikas Dhiman
arXiv Computer Vision
Sep 17

CADSplat: Sparse-View 3D Gaussian Splatting Aided by CAD Models for Robust, Photorealistic Digital-Twin Reconstruction

CADSplat is a framework that reconstructs photorealistic, geometrically accurate digital twins from fewer than 15 wide‑baseline images by regularizing 3D Gaussian Splatting with an explicit CAD shape prior. It matches segmented object silhouettes to a CAD library to retrieve a suitable model and camera poses, then anchors Gaussian primitives to the model’s surface and jointly optimizes splat parameters, registration, and a non‑rigid deformation field. Experiments on two real‑world datasets show CADSplat outperforms baselines, especially in sparse and self‑occluded scenarios, and its gains mainly stem from constraining splats to a surface rather than the CAD shape itself.

By Kristof Overdulve, Lode Jorissen, Nick Michiels