Hugging Face Trending Papers

ManiSplat: Manipulation Trajectory Synthesis from Monocular Video via Decoupled 3D Gaussian Splatting

Reconstructing dynamic and interactive 3D scenes from real-world observations remains a fundamental challenge in computer vision and robotics. While recent advances in 3D Gaussian Splatting have enabled high-fidelity static reconstruction, extending it to interactive environments with articulated robots and manipulable objects remains difficult due to complex contact interactions and abrupt pose changes.

arXiv Computer Vision
Sep 24

Track2Art: Motion-Centric Articulated Object Model Recovery from 2D Point Trackers

Track2Art is a motion‑centric framework that recovers articulated object models from RGB‑D interaction videos by lifting 2D point tracks into 3D trajectories. It groups these trajectories into rigid‑part hypotheses and uses learned‑analytic reasoning to infer directed kinematic relations, joint types, and joint geometry. On the PartNet‑Mobility benchmark, it achieves 0.695 Point IoU and 0.410 end‑to‑end J@20 without requiring ground‑truth part counts or test‑time optimization.

By Xiaotong Li, Yixiong Jing, Junsheng Ding, Weihang Li, Benjamin Busam, Guangming Wang, Brian Sheil
arXiv Computer Vision
Aug 28

Grasp in Gaussians: Fast Monocular Reconstruction of Dynamic Hand-Object Interactions

Grasp in Gaussians (GraG) is a fast, robust method for reconstructing dynamic 3D hand‑object interactions from a single monocular video. It leverages pretrained hand and object priors and represents the scene with a compact Sum‑of‑Gaussians (SoG) model, enabling efficient tracking while preserving geometric fidelity. Experiments show GraG achieves temporally coherent reconstructions on long sequences 4.4×–38.9× faster than prior work.

By Ayce Idil Aytekin, Xu Chen, Zhengyang Shen, Thabo Beeler, Helge Rhodin, Rishabh Dabral, Christian Theobalt
arXiv AI
Sep 24

InfiNoVA: Infinite Novel View Augmentation for Viewpoint Invariant Robot Policies

InfiNoVA is a data‑augmentation framework that transforms synchronized multi‑camera demonstrations into a dense, geometrically consistent set of training views by reconstructing each manipulation trajectory as a time‑varying 3D Gaussian. The method renders novel observations from sampled camera poses while preserving the original state‑action pairs, improving frame‑level fidelity and temporal consistency compared to generative synthesis. Across four real‑world manipulation tasks, policies trained with InfiNoVA achieve 5.4× higher average success under unseen randomized viewpoints than VISTA‑based augmentation and 1.7× higher success than training on all five physical camera views.

By Sai Puneeth Reddy Gottam, Elmar Rueckert, Vedant Dave