Track2Art is a motion‑centric framework that recovers articulated object models from RGB‑D interaction videos by lifting 2D point tracks into 3D trajectories. It groups these trajectories into rigid‑part hypotheses and uses learned‑analytic reasoning to infer directed kinematic relations, joint types, and joint geometry. On the PartNet‑Mobility benchmark, it achieves 0.695 Point IoU and 0.410 end‑to‑end J@20 without requiring ground‑truth part counts or test‑time optimization.
By Xiaotong Li, Yixiong Jing, Junsheng Ding, Weihang Li, Benjamin Busam, Guangming Wang, Brian Sheil
arXiv:2609.19119v1 Announce Type: new
Abstract: Human videos contain rich causal evidence for robot manipulation: they reveal how hand motion induces object motion and produces task-relevant changes...
By Jiaming Zhang, Homanga Bharadhwaj
Artic-O is an end‑to‑end, feed‑forward framework that reconstructs articulated objects from sparse images by learning latent geometry. It maps multi‑state observations into a pretrained latent geometry space, uses a frozen flow‑matching decoder for complete‑shape priors, and fuses visual tokens with geometry latents in an image‑grounded part‑reasoning module to segment active parts and predict articulation. Trained with a geometry‑to‑articulation curriculum and a decoupled two‑pass strategy, Artic‑O achieves high reconstruction quality and articulation accuracy while drastically reducing inference time from 9 minutes to about 0.3 seconds per object.
By Xuyang Wang, Zhenyu Li, Jian Ding, Habib Slim, Peter Wonka, Hongdong Li, Mohamed Elhoseiny
World models endow perceptual systems with the ability to predict how scenes evolve under interaction. They are most beneficial when trained on diverse volumes of data, to instill a rich prior into do...
arXiv:2609.19142v1 Announce Type: new
Abstract: World models endow perceptual systems with the ability to predict how scenes evolve under interaction. They are most beneficial when trained on diverse...
By Bardienus P. Duisterhof, Kaifeng Zhang, Adam Hung, Bowen Wen, Stan Birchfield, Yunzhu Li, Deva Ramanan, Jeffrey Ichnowski
arXiv:2605.14854v3 Announce Type: replace-cross
Abstract: Human Mesh Recovery (HMR) is fundamentally ambiguous: under occlusion or weak depth cues, multiple 3D bodies can explain the same image evide...
By Patrick Kwon, Chen Chen
arXiv:2509.04276v3 Announce Type: replace
Abstract: We present a method for modeling articulated objects from sparse images with unknown camera poses. Existing approaches require dense multi-view obs...
By Jianning Deng, Kartic Subr, Hakan Bilen
Point2Pose is a model‑free method for causal 6D pose tracking of multiple rigid objects using monocular RGB‑D video. It starts from sparse image points and employs a 2D point tracker to maintain long‑range correspondences, allowing instant recovery after complete occlusion. The system also incrementally builds an online Truncated Signed Distance Function (TSDF) representation of the tracked objects and introduces a new multi‑object tracking dataset with motion‑capture ground truth.
By Tzu-Yuan Lin, Ho Jae Lee, Kevin Doherty, Yonghyeon Lee, Sangbae Kim
arXiv:2609.24487v1 Announce Type: new
Abstract: In this work, we present a method for shape reconstruction and tracking from video via agentic analysis-by-synthesis. Unlike prior methods which first...
By Kirill Mazur, Nikita Karaev, Matthew Chang, Jitendra Malik, Nur Muhammad "Mahi'' Shafiullah
Reconstructing dynamic and interactive 3D scenes from real-world observations remains a fundamental challenge in computer vision and robotics. While recent advances in 3D Gaussian Splatting have enabled high-fidelity static reconstruction, extending it to interactive environments with articulated robots and manipulable objects remains difficult due to complex contact interactions and abrupt pose changes.
FunArt is a framework that builds articulation‑aware functional 3D scene graphs from a single static RGB‑D observation. It reconstructs object instances, converts their geometry into the O‑Voxel representation of TRELLIS.2, and uses a frozen sparse‑compression VAE as a structural prior. A lightweight query‑based decoder jointly segments movable parts and functional interactive elements while estimating motion type, axis, origin, and range, achieving state‑of‑the‑art results on the Articulate3D dataset.
By Dennis Rotondi, Abdelrhman Werby, Kai O. Arras
The key challenge in articulated 3D object generation from a single image is accurately predicting the underlying kinematic structure. Existing methods either infer kinematic parameters directly from a static image that lacks dynamic part-level kinematic relationships, or estimate parameters from visual dynamics generated from a single image, which is prone to accumulated errors of two steps.