arXiv AI

FACT: Fidelity-Aware Construction of Articulated Twins

arXiv Computer Vision
Sep 24

Track2Art: Motion-Centric Articulated Object Model Recovery from 2D Point Trackers

Track2Art is a motion‑centric framework that recovers articulated object models from RGB‑D interaction videos by lifting 2D point tracks into 3D trajectories. It groups these trajectories into rigid‑part hypotheses and uses learned‑analytic reasoning to infer directed kinematic relations, joint types, and joint geometry. On the PartNet‑Mobility benchmark, it achieves 0.695 Point IoU and 0.410 end‑to‑end J@20 without requiring ground‑truth part counts or test‑time optimization.

By Xiaotong Li, Yixiong Jing, Junsheng Ding, Weihang Li, Benjamin Busam, Guangming Wang, Brian Sheil
Hugging Face Trending Papers
Jun 9

ManiSplat: Manipulation Trajectory Synthesis from Monocular Video via Decoupled 3D Gaussian Splatting

Reconstructing dynamic and interactive 3D scenes from real-world observations remains a fundamental challenge in computer vision and robotics. While recent advances in 3D Gaussian Splatting have enabled high-fidelity static reconstruction, extending it to interactive environments with articulated robots and manipulable objects remains difficult due to complex contact interactions and abrupt pose changes.

arXiv Computer Vision
Sep 11

Artic-O: End-to-End Articulated Object Reconstruction via Latent Geometry Learning

Artic-O is an end‑to‑end, feed‑forward framework that reconstructs articulated objects from sparse images by learning latent geometry. It maps multi‑state observations into a pretrained latent geometry space, uses a frozen flow‑matching decoder for complete‑shape priors, and fuses visual tokens with geometry latents in an image‑grounded part‑reasoning module to segment active parts and predict articulation. Trained with a geometry‑to‑articulation curriculum and a decoupled two‑pass strategy, Artic‑O achieves high reconstruction quality and articulation accuracy while drastically reducing inference time from 9 minutes to about 0.3 seconds per object.

By Xuyang Wang, Zhenyu Li, Jian Ding, Habib Slim, Peter Wonka, Hongdong Li, Mohamed Elhoseiny
arXiv Computer Vision
Aug 28

Reconstructing Humans and Objects in Interaction using Large Reconstruction Models

The paper introduces MILO, a framework that uses Large Reconstruction Models (LRMs) to reconstruct detailed 3D human‑object interactions from a single image. By treating the LRM mesh as a geometric scaffold, MILO segments it into human and object parts, fits a parametric body model to the human component, and optionally aligns an object template to the object component. The approach achieves higher reconstruction accuracy than existing baselines across multiple benchmarks and interaction scenarios.

By Agniv Chatterjee, Georgios Pavlakos