Interact3D: Compositional 3D Generation of Interactive Objects
arXiv:2603. 16085v2 Announce Type: replace-cross Abstract: Recent breakthroughs in 3D generation have enabled the synthesis of high-fidelity individual assets.
The paper presents a method for reconstructing object geometry under occlusion by combining generative shape priors with contact-based constraints. Generative models provide plausible guesses for unseen parts, while contact information from videos or physical interactions supplies sparse boundary constraints. The authors integrate these cues through a contact-guided 3D generation framework, demonstrating improved reconstruction on synthetic and real-world datasets compared to baseline approaches.
arXiv:2603. 16085v2 Announce Type: replace-cross Abstract: Recent breakthroughs in 3D generation have enabled the synthesis of high-fidelity individual assets.
arXiv:2602. 08058v3 Announce Type: replace-cross Abstract: In the presence of occlusions and measurement noise, geometrically accurate scene reconstructions -- which fit the sensor data -- can still be physically incorrect.
The paper introduces MILO, a framework that uses Large Reconstruction Models (LRMs) to reconstruct detailed 3D human‑object interactions from a single image. By treating the LRM mesh as a geometric scaffold, MILO segments it into human and object parts, fits a parametric body model to the human component, and optionally aligns an object template to the object component. The approach achieves higher reconstruction accuracy than existing baselines across multiple benchmarks and interaction scenarios.
arXiv:2608.21776v1 Announce Type: new Abstract: Recent advances in image editing allow impressive manipulation of objects, existing methods still struggle to handle spatial movement in complex scenes...
arXiv:2510. 11014v2 Announce Type: replace-cross Abstract: Autonomous robots often view rooms only partially, through a doorway, where the walls and scene structure hide the geometry and task-relevant semantics needed for safe navigation and goal-directed action.
World models offer a promising route toward robot planning by enabling agents to imagine and verify the consequences of actions before execution. However, current video-based world models often struggle to capture the physical constraints that govern manipulation, particularly contact.
Reconstructing articulated objects with multiple movable parts is essential for understanding object structure and enabling physical interaction. However, this reconstruction task poses significant challenges due to the entanglement of geometry, appearance, and motion parameters during optimization.
arXiv:2606. 24450v1 Announce Type: cross Abstract: Perceiving physical contact is fundamental to dexterous manipulation.
arXiv:2509. 06191v2 Announce Type: replace-cross Abstract: Recent 3D generative models, which are capable of generating full object shapes from just a few images, now open up new opportunities in robotics.
arXiv:2606. 08057v1 Announce Type: cross Abstract: Egocentric RGB-D videos offer a natural source of human dexterous manipulation demonstrations, but existing data is difficult to use for robot learning because object pose, geometry, and contact information are often missing or require pre-scanned object assets.
arXiv:2608.28386v1 Announce Type: new Abstract: Existing monocular full-body 3D human-object interaction (HOI) methods do not combine explicit finger-level grasp optimization with category-agnostic o...
arXiv:2606. 28813v1 Announce Type: cross Abstract: Human videos are a scalable source of supervision for robot manipulation, as they are abundant and naturally capture rich object interactions.