arXiv AI

Human4K: A Large-Scale 4K Multi-View Mocap Dataset for Whole-Body 3D Human Reconstruction

arXiv:2607. 13646v1 Announce Type: cross Abstract: Recent advances in 3D human reconstruction have improved overall performance, yet current models still fail in the most challenging real-world scenarios.

arXiv Computer Vision
Sep 25

Ego-Exo4D Human Meshes Dataset: 4D Human Motion Reconstruction for Ego-Exo Captures

Ego-Exo4D is a large-scale dataset that offers synchronized egocentric and multi-view exocentric video for applications such as skill learning, procedural activity understanding, and embodied AI. The original dataset only includes sparse 3D human pose annotations, making dense motion reconstruction challenging. To address this, the authors introduce Ego-Exo4D-HM, a new dataset containing 4D human motion reconstructions for the Ego-Exo4D captures, along with a reconstruction pipeline and accompanying code and documentation.

By Abhiram Maddukuri, Georgios Pavlakos
arXiv Computer Vision
Sep 2

Feed-Forward Multi-view Multi-person Reconstruction with Contrastive Human-Aware 3D Representation

The paper introduces a top‑down approach for multi‑person 3D reconstruction from multiple views, using a unified, instance‑centric human‑aware 3D space. Observations from different cameras are lifted into this shared space where geometry, appearance, and semantic cues are jointly encoded, and a spatial contrastive learning strategy aligns features of the same person across views while separating different individuals. The method then regresses SMPL parameters from 3D tokens in a feed‑forward manner, achieving robust, accurate, and efficient reconstruction even under severe occlusions.

By Yuanwang Yang, Buzhen Huang, Zongxuan Ren, Jing Huang, Kun Li
arXiv AI
Jun 29

HAT-4D: Lifting Monocular Video for 4D Multi-Object Interactions via Human-Agent Collaboration

arXiv:2606. 28215v1 Announce Type: cross Abstract: Extracting dynamic 4D object interactions from massive, in-the-wild monocular videos offers a highly efficient data collection pathway for scaling Embodied AI and training VLAs.

By Jiaxin Li, Yuxiang Wu, Zhenkai Zhang, Xinrui Shi, Haoyuan Wang, Yichen Zhao, Su Linxiang, Chenyang Yu, Mingyu Zhang, Yifan Ding, Boran Wen, Li Zhang, Ruiyang Liu, Yong-Lu Li
arXiv Computer Vision
Sep 18

DirtyMoCap: Robust Motion Capture from Unconstrained Markers

DirtyMoCap is a marker‑layout‑free framework that converts unordered, noisy optical motion capture markers into a fixed set of proxy anchors representing skeletal joints and body surface points. Using a recurrent sliding‑window architecture to track these anchors and a custom differentiable Gauss‑Newton solver to fit the SMPL‑H model, the method learns adaptive observation confidence, smoothness, and prior weights end‑to‑end. Experiments show that DirtyMoCap generalizes across arbitrary marker configurations, outperforms configuration‑specific baselines in joint and vertex accuracy, and achieves up to a 100× speedup over standard PyTorch implementations, enabling the creation of a temporally coherent Kung Fu motion dataset.

By Long Wang, Shuting Zhao, Shen Yan, Siyuan Yu, Xiaoben Li, Zeyu Cai, Yumeng Hou, Yuliang Xiu