arXiv AI By Tianshun Han, Ziyu Shi, Lijian Liu, Ajian Liu, Benjia Zhou, Hugo Jair Escalante, Yanyan Liang, Sergio Escalera, Zhen Lei, Jun Wan

Human4K: A Large-Scale 4K Multi-View Mocap Dataset for Whole-Body 3D Human Reconstruction

Read the original on arXiv AI →

arXiv:2607. 13646v1 Announce Type: cross Abstract: Recent advances in 3D human reconstruction have improved overall performance, yet current models still fail in the most challenging real-world scenarios.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Sep 25

Ego-Exo4D Human Meshes Dataset: 4D Human Motion Reconstruction for Ego-Exo Captures

Ego-Exo4D is a large-scale dataset that offers synchronized egocentric and multi-view exocentric video for applications such as skill learning, procedural activity understanding, and embodied AI. The original dataset only includes sparse 3D human pose annotations, making dense motion reconstruction challenging. To address this, the authors introduce Ego-Exo4D-HM, a new dataset containing 4D human motion reconstructions for the Ego-Exo4D captures, along with a reconstruction pipeline and accompanying code and documentation.

By Abhiram Maddukuri, Georgios Pavlakos
arXiv Computer Vision
Sep 2

Feed-Forward Multi-view Multi-person Reconstruction with Contrastive Human-Aware 3D Representation

The paper introduces a top‑down approach for multi‑person 3D reconstruction from multiple views, using a unified, instance‑centric human‑aware 3D space. Observations from different cameras are lifted into this shared space where geometry, appearance, and semantic cues are jointly encoded, and a spatial contrastive learning strategy aligns features of the same person across views while separating different individuals. The method then regresses SMPL parameters from 3D tokens in a feed‑forward manner, achieving robust, accurate, and efficient reconstruction even under severe occlusions.

By Yuanwang Yang, Buzhen Huang, Zongxuan Ren, Jing Huang, Kun Li
arXiv AI
Jun 29

HAT-4D: Lifting Monocular Video for 4D Multi-Object Interactions via Human-Agent Collaboration

arXiv:2606. 28215v1 Announce Type: cross Abstract: Extracting dynamic 4D object interactions from massive, in-the-wild monocular videos offers a highly efficient data collection pathway for scaling Embodied AI and training VLAs.

By Jiaxin Li, Yuxiang Wu, Zhenkai Zhang, Xinrui Shi, Haoyuan Wang, Yichen Zhao, Su Linxiang, Chenyang Yu, Mingyu Zhang, Yifan Ding, Boran Wen, Li Zhang, Ruiyang Liu, Yong-Lu Li