arXiv AI

FactorizedHMR: A Hybrid Framework for Video Human Mesh Recovery

arXiv Computer Vision
Sep 11

Artic-O: End-to-End Articulated Object Reconstruction via Latent Geometry Learning

Artic-O is an end‑to‑end, feed‑forward framework that reconstructs articulated objects from sparse images by learning latent geometry. It maps multi‑state observations into a pretrained latent geometry space, uses a frozen flow‑matching decoder for complete‑shape priors, and fuses visual tokens with geometry latents in an image‑grounded part‑reasoning module to segment active parts and predict articulation. Trained with a geometry‑to‑articulation curriculum and a decoupled two‑pass strategy, Artic‑O achieves high reconstruction quality and articulation accuracy while drastically reducing inference time from 9 minutes to about 0.3 seconds per object.

By Xuyang Wang, Zhenyu Li, Jian Ding, Habib Slim, Peter Wonka, Hongdong Li, Mohamed Elhoseiny
arXiv Computer Vision
Sep 24

Track2Art: Motion-Centric Articulated Object Model Recovery from 2D Point Trackers

Track2Art is a motion‑centric framework that recovers articulated object models from RGB‑D interaction videos by lifting 2D point tracks into 3D trajectories. It groups these trajectories into rigid‑part hypotheses and uses learned‑analytic reasoning to infer directed kinematic relations, joint types, and joint geometry. On the PartNet‑Mobility benchmark, it achieves 0.695 Point IoU and 0.410 end‑to‑end J@20 without requiring ground‑truth part counts or test‑time optimization.

By Xiaotong Li, Yixiong Jing, Junsheng Ding, Weihang Li, Benjamin Busam, Guangming Wang, Brian Sheil
arXiv Computer Vision
Sep 15

MoCapAnything V2: End-to-End Motion Capture for Arbitrary Skeletons

arXiv:2604.28130v4 Announce Type: replace Abstract: Recent methods for arbitrary-skeleton motion capture from monocular video follow a factorized pipeline, where a Video-to-Pose network predicts join...

By Kehong Gong, Zhengyu Wen, Dao Thien Phong, Mingxi Xu, Weixia He, Qi Wang, Ning Zhang, Zhengyu Li, Guanli Hou, Dongze Lian, Xiaoyu He, Mingyuan Zhang, Hanwang Zhang
arXiv AI
6d ago

WildHSR: Metric Feed-Forward 4D People-Scene Reconstruction from a 3D Foundation Model

WildHSR introduces a lightweight adaptation of 3D foundation models to jointly recover metric cameras, scene geometry, and persistent person identities from monocular video. By generating pseudo‑scale labels from curated web footage and fine‑tuning a Scale Readout, the method predicts metric scale directly from foundation‑model tokens. It also exploits intermediate query‑key features to associate per‑frame bodies, enabling feed‑forward reconstruction that outperforms state‑of‑the‑art optimization‑based methods on several benchmarks while running at 10.1 fps.

By Jerrin Bright, John Zelek
arXiv Computer Vision
Sep 18

DirtyMoCap: Robust Motion Capture from Unconstrained Markers

DirtyMoCap is a marker‑layout‑free framework that converts unordered, noisy optical motion capture markers into a fixed set of proxy anchors representing skeletal joints and body surface points. Using a recurrent sliding‑window architecture to track these anchors and a custom differentiable Gauss‑Newton solver to fit the SMPL‑H model, the method learns adaptive observation confidence, smoothness, and prior weights end‑to‑end. Experiments show that DirtyMoCap generalizes across arbitrary marker configurations, outperforms configuration‑specific baselines in joint and vertex accuracy, and achieves up to a 100× speedup over standard PyTorch implementations, enabling the creation of a temporally coherent Kung Fu motion dataset.

By Long Wang, Shuting Zhao, Shen Yan, Siyuan Yu, Xiaoben Li, Zeyu Cai, Yumeng Hou, Yuliang Xiu