arXiv Computer Vision By Agniv Chatterjee, Georgios Pavlakos

Reconstructing Humans and Objects in Interaction using Large Reconstruction Models

Read the original on arXiv Computer Vision →

The paper introduces MILO, a framework that uses Large Reconstruction Models (LRMs) to reconstruct detailed 3D human‑object interactions from a single image. By treating the LRM mesh as a geometric scaffold, MILO segments it into human and object parts, fits a parametric body model to the human component, and optionally aligns an object template to the object component. The approach achieves higher reconstruction accuracy than existing baselines across multiple benchmarks and interaction scenarios.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv AI
Jun 4

SAM 3D: 3Dfy Anything in Images

arXiv:2511. 16624v2 Announce Type: replace-cross Abstract: We present SAM 3D, a generative model for visually grounded 3D object reconstruction, predicting geometry, texture, and layout from a single image.

By SAM 3D Team, Xingyu Chen, Fu-Jen Chu, Pierre Gleize, Kevin J Liang, Alexander Sax, Hao Tang, Weiyao Wang, Michelle Guo, Thibaut Hardin, Xiang Li, Aohan Lin, Jiawei Liu, Ziqi Ma, Anushka Sagar, Bowen Song, Xiaodong Wang, Jianing Yang, Bowen Zhang, Piotr Doll\'ar, Georgia Gkioxari, Matt Feiszli, Jitendra Malik
arXiv AI
Jun 29

HAT-4D: Lifting Monocular Video for 4D Multi-Object Interactions via Human-Agent Collaboration

arXiv:2606. 28215v1 Announce Type: cross Abstract: Extracting dynamic 4D object interactions from massive, in-the-wild monocular videos offers a highly efficient data collection pathway for scaling Embodied AI and training VLAs.

By Jiaxin Li, Yuxiang Wu, Zhenkai Zhang, Xinrui Shi, Haoyuan Wang, Yichen Zhao, Su Linxiang, Chenyang Yu, Mingyu Zhang, Yifan Ding, Boran Wen, Li Zhang, Ruiyang Liu, Yong-Lu Li
arXiv Computer Vision
1d ago

Feed-Forward Multi-view Multi-person Reconstruction with Contrastive Human-Aware 3D Representation

The paper introduces a top‑down approach for multi‑person 3D reconstruction from multiple views, using a unified, instance‑centric human‑aware 3D space. Observations from different cameras are lifted into this shared space where geometry, appearance, and semantic cues are jointly encoded, and a spatial contrastive learning strategy aligns features of the same person across views while separating different individuals. The method then regresses SMPL parameters from 3D tokens in a feed‑forward manner, achieving robust, accurate, and efficient reconstruction even under severe occlusions.

By Yuanwang Yang, Buzhen Huang, Zongxuan Ren, Jing Huang, Kun Li