arXiv Computer Vision

Reconstructing Humans and Objects in Interaction using Large Reconstruction Models

The paper introduces MILO, a framework that uses Large Reconstruction Models (LRMs) to reconstruct detailed 3D human‑object interactions from a single image. By treating the LRM mesh as a geometric scaffold, MILO segments it into human and object parts, fits a parametric body model to the human component, and optionally aligns an object template to the object component. The approach achieves higher reconstruction accuracy than existing baselines across multiple benchmarks and interaction scenarios.

arXiv AI
Jun 4

SAM 3D: 3Dfy Anything in Images

arXiv:2511. 16624v2 Announce Type: replace-cross Abstract: We present SAM 3D, a generative model for visually grounded 3D object reconstruction, predicting geometry, texture, and layout from a single image.

By SAM 3D Team, Xingyu Chen, Fu-Jen Chu, Pierre Gleize, Kevin J Liang, Alexander Sax, Hao Tang, Weiyao Wang, Michelle Guo, Thibaut Hardin, Xiang Li, Aohan Lin, Jiawei Liu, Ziqi Ma, Anushka Sagar, Bowen Song, Xiaodong Wang, Jianing Yang, Bowen Zhang, Piotr Doll\'ar, Georgia Gkioxari, Matt Feiszli, Jitendra Malik
arXiv AI
Jun 29

HAT-4D: Lifting Monocular Video for 4D Multi-Object Interactions via Human-Agent Collaboration

arXiv:2606. 28215v1 Announce Type: cross Abstract: Extracting dynamic 4D object interactions from massive, in-the-wild monocular videos offers a highly efficient data collection pathway for scaling Embodied AI and training VLAs.

By Jiaxin Li, Yuxiang Wu, Zhenkai Zhang, Xinrui Shi, Haoyuan Wang, Yichen Zhao, Su Linxiang, Chenyang Yu, Mingyu Zhang, Yifan Ding, Boran Wen, Li Zhang, Ruiyang Liu, Yong-Lu Li
arXiv Computer Vision
1d ago

Feed-Forward Multi-view Multi-person Reconstruction with Contrastive Human-Aware 3D Representation

The paper introduces a top‑down approach for multi‑person 3D reconstruction from multiple views, using a unified, instance‑centric human‑aware 3D space. Observations from different cameras are lifted into this shared space where geometry, appearance, and semantic cues are jointly encoded, and a spatial contrastive learning strategy aligns features of the same person across views while separating different individuals. The method then regresses SMPL parameters from 3D tokens in a feed‑forward manner, achieving robust, accurate, and efficient reconstruction even under severe occlusions.

By Yuanwang Yang, Buzhen Huang, Zongxuan Ren, Jing Huang, Kun Li
arXiv Computer Vision
Aug 26

SceneReGen: Generative Reconstruction of 3D Scenes from a Single Image

SceneReGen is a new framework for reconstructing 3D scenes from a single image by generating and assembling complete object meshes within a shared observation‑aligned scene frame. It uses selective pose factorization to encode each object’s observed orientation directly into the generated mesh, while estimating translation and scale from instance‑level and global scene cues. Evaluated on the 3D‑FUTURE dataset, SceneReGen outperforms existing methods on scene‑level metrics and shows strong performance on object‑level metrics, demonstrating its effectiveness in autonomous‑driving and embodied‑AI scenarios.

By Zefan Tian, Yuteng Ye, Yiheng Zhang, Yuhang Yang, Xueqiang Lv, Shizhou Zhang, Le Liu, Di Xu