Artic-O is an end‑to‑end, feed‑forward framework that reconstructs articulated objects from sparse images by learning latent geometry. It maps multi‑state observations into a pretrained latent geometry space, uses a frozen flow‑matching decoder for complete‑shape priors, and fuses visual tokens with geometry latents in an image‑grounded part‑reasoning module to segment active parts and predict articulation. Trained with a geometry‑to‑articulation curriculum and a decoupled two‑pass strategy, Artic‑O achieves high reconstruction quality and articulation accuracy while drastically reducing inference time from 9 minutes to about 0.3 seconds per object.
By Xuyang Wang, Zhenyu Li, Jian Ding, Habib Slim, Peter Wonka, Hongdong Li, Mohamed Elhoseiny
arXiv:2609.27675v2 Announce Type: replace
Abstract: Understanding articulated objects is fundamental for robotic interaction, requiring accurate rigid-part discovery and the recovery of their kinemat...
By Xiaotong Li, Yixiong Jing, Junsheng Ding, Weihang Li, Benjamin Busam, Guangming Wang, Brian Sheil
Track2Art is a motion‑centric framework that recovers articulated object models from RGB‑D interaction videos by lifting 2D point tracks into 3D trajectories. It groups these trajectories into rigid‑part hypotheses and uses learned‑analytic reasoning to infer directed kinematic relations, joint types, and joint geometry. On the PartNet‑Mobility benchmark, it achieves 0.695 Point IoU and 0.410 end‑to‑end J@20 without requiring ground‑truth part counts or test‑time optimization.
By Xiaotong Li, Yixiong Jing, Junsheng Ding, Weihang Li, Benjamin Busam, Guangming Wang, Brian Sheil
arXiv:2509.04276v3 Announce Type: replace
Abstract: We present a method for modeling articulated objects from sparse images with unknown camera poses. Existing approaches require dense multi-view obs...
By Jianning Deng, Kartic Subr, Hakan Bilen
arXiv:2609.01276v1 Announce Type: new
Abstract: Complete 3D perception from egocentric video requires recovering the surrounding scene and the wearer's full-body motion in a shared metric frame. Exis...
By Kai Guan, Minchao Jiang, Ruichen WangLi, Wentao Zhu, Lei Zhang
Human mesh recovery (HMR) aims to recover 3D human meshes from images. Most existing HMR benchmarks and methods focus on either multi-person reconstruction from a single view or single-person reconstruction from multiple views, where the number of subjects and the scene scale are relatively limited.
arXiv:2603.12789v3 Announce Type: replace
Abstract: Recent advances in 3D foundation models have enabled joint reconstruction of humans and their surrounding environments. However, combining independ...
By Sangmin Kim, Minhyuk Hwang, Geonho Cha, Dongyoon Wee, Jaesik Park
arXiv:2604.28130v4 Announce Type: replace
Abstract: Recent methods for arbitrary-skeleton motion capture from monocular video follow a factorized pipeline, where a Video-to-Pose network predicts join...
By Kehong Gong, Zhengyu Wen, Dao Thien Phong, Mingxi Xu, Weixia He, Qi Wang, Ning Zhang, Zhengyu Li, Guanli Hou, Dongze Lian, Xiaoyu He, Mingyuan Zhang, Hanwang Zhang
arXiv:2609.36937v1 Announce Type: cross
Abstract: Human image animation aims to transfer motion from a driving video to subjects in a reference image. Despite remarkable progress in video generation,...
By Sangeyl Lee, Seunghyun Shin, Seungho Park, Wooseok Jeon, Hae-Gon Jeon
WildHSR introduces a lightweight adaptation of 3D foundation models to jointly recover metric cameras, scene geometry, and persistent person identities from monocular video. By generating pseudo‑scale labels from curated web footage and fine‑tuning a Scale Readout, the method predicts metric scale directly from foundation‑model tokens. It also exploits intermediate query‑key features to associate per‑frame bodies, enabling feed‑forward reconstruction that outperforms state‑of‑the‑art optimization‑based methods on several benchmarks while running at 10.1 fps.
By Jerrin Bright, John Zelek
DirtyMoCap is a marker‑layout‑free framework that converts unordered, noisy optical motion capture markers into a fixed set of proxy anchors representing skeletal joints and body surface points. Using a recurrent sliding‑window architecture to track these anchors and a custom differentiable Gauss‑Newton solver to fit the SMPL‑H model, the method learns adaptive observation confidence, smoothness, and prior weights end‑to‑end. Experiments show that DirtyMoCap generalizes across arbitrary marker configurations, outperforms configuration‑specific baselines in joint and vertex accuracy, and achieves up to a 100× speedup over standard PyTorch implementations, enabling the creation of a temporally coherent Kung Fu motion dataset.
By Long Wang, Shuting Zhao, Shen Yan, Siyuan Yu, Xiaoben Li, Zeyu Cai, Yumeng Hou, Yuliang Xiu
arXiv:2609.19119v1 Announce Type: new
Abstract: Human videos contain rich causal evidence for robot manipulation: they reveal how hand motion induces object motion and produces task-relevant changes...
By Jiaming Zhang, Homanga Bharadhwaj