arXiv Computer Vision

CrowdOcc: Monocular Semantic Scene Completion for Quadruped Robots in Crowded Indoor Environments

CrowdOcc introduces a new RGB‑D dataset and monocular semantic scene completion framework tailored for quadruped robots navigating crowded indoor spaces. The dataset comprises 25.1K frames across 11 scenes, with annotations derived via static‑dynamic decoupling. The framework employs Normal Guided Scene Geometry Fusion and Human‑Centric Sparse Interaction to robustly model occluded geometry and human‑scene relationships, achieving state‑of‑the‑art performance on the dataset’s test set.

arXiv Computer Vision
Sep 3

OccAnyScene: Towards Unified Indoor-Outdoor 3D Occupancy Prediction

OccAnyScene introduces a unified approach for 3D occupancy prediction across both indoor and outdoor scenes, addressing the challenge of handling diverse camera setups, spatial ranges, voxel specifications, and semantic taxonomies. The method builds on a pretrained depth model, using pixel-aligned frustum feature aggregation and frustum-parameterized Gaussian construction to generate scene-adaptive occupancy predictions. It achieves state-of-the-art performance, scoring 59.92% mIoU on Occ-ScanNet and 23.06% mIoU on SurroundOcc-nuScenes.

By Junjie Liu, Wanshui Gan, Zitong Dai, Guiping Cao, Yan Li, Ke Chen, Dongmei Jiang, Jianguo Zhang, Xiangyuan Lan
arXiv AI
Sep 24

MessyKitchens: Contact-rich object-level 3D scene reconstruction

MessyKitchens introduces a new dataset of cluttered real-world kitchen scenes with detailed 3D object shapes, poses, and accurate contact information. The authors extend the SAM 3D single-object reconstruction method with a Multi-Object Decoder (MOD) to jointly reconstruct entire scenes, achieving better registration accuracy and reduced inter-object penetration compared to prior work. The dataset, benchmark, code, and pretrained models will be publicly released on the project website.

By Junaid Ahmed Ansari, Ran Ding, Fabio Pizzati, Ivan Laptev
Hugging Face Trending Papers
Jun 22

Humanoid-OmniOcc: Stereo-Based Full-View Occupancy Dataset for Embodied AI

Occupancy prediction at voxel-level granularity is essential for safe robotic navigation and interaction in complex environments. Existing occupancy datasets, however, are predominantly designed for autonomous driving with vehicle-centric biases -- forward-facing cameras, far-field geometry, and static road priors -- limiting their applicability to embodied humanoid perception.

arXiv AI
Jun 29

HAT-4D: Lifting Monocular Video for 4D Multi-Object Interactions via Human-Agent Collaboration

arXiv:2606. 28215v1 Announce Type: cross Abstract: Extracting dynamic 4D object interactions from massive, in-the-wild monocular videos offers a highly efficient data collection pathway for scaling Embodied AI and training VLAs.

By Jiaxin Li, Yuxiang Wu, Zhenkai Zhang, Xinrui Shi, Haoyuan Wang, Yichen Zhao, Su Linxiang, Chenyang Yu, Mingyu Zhang, Yifan Ding, Boran Wen, Li Zhang, Ruiyang Liu, Yong-Lu Li
arXiv AI
Jun 8

MatterDoor: Sampling Zero-shot Spatio-semantic Priors using Generative Models

arXiv:2510. 11014v2 Announce Type: replace-cross Abstract: Autonomous robots often view rooms only partially, through a doorway, where the walls and scene structure hide the geometry and task-relevant semantics needed for safe navigation and goal-directed action.

By Subhransu S. Bhattacharjee, Hao Lu, Dylan Campbell, Rahul Shome