arXiv Computer Vision By Feiyang Chen, Jincheng Hu, Yiduo Chen, Jihao Li, Yue Liang, Bingzhao Gao, Yanjun Huang, Yuanjian Zhang

CrowdOcc: Monocular Semantic Scene Completion for Quadruped Robots in Crowded Indoor Environments

Read the original on arXiv Computer Vision →

CrowdOcc introduces a new RGB‑D dataset and monocular semantic scene completion framework tailored for quadruped robots navigating crowded indoor spaces. The dataset comprises 25.1K frames across 11 scenes, with annotations derived via static‑dynamic decoupling. The framework employs Normal Guided Scene Geometry Fusion and Human‑Centric Sparse Interaction to robustly model occluded geometry and human‑scene relationships, achieving state‑of‑the‑art performance on the dataset’s test set.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Computer Vision
Sep 3

OccAnyScene: Towards Unified Indoor-Outdoor 3D Occupancy Prediction

OccAnyScene introduces a unified approach for 3D occupancy prediction across both indoor and outdoor scenes, addressing the challenge of handling diverse camera setups, spatial ranges, voxel specifications, and semantic taxonomies. The method builds on a pretrained depth model, using pixel-aligned frustum feature aggregation and frustum-parameterized Gaussian construction to generate scene-adaptive occupancy predictions. It achieves state-of-the-art performance, scoring 59.92% mIoU on Occ-ScanNet and 23.06% mIoU on SurroundOcc-nuScenes.

By Junjie Liu, Wanshui Gan, Zitong Dai, Guiping Cao, Yan Li, Ke Chen, Dongmei Jiang, Jianguo Zhang, Xiangyuan Lan
arXiv AI
Sep 24

MessyKitchens: Contact-rich object-level 3D scene reconstruction

MessyKitchens introduces a new dataset of cluttered real-world kitchen scenes with detailed 3D object shapes, poses, and accurate contact information. The authors extend the SAM 3D single-object reconstruction method with a Multi-Object Decoder (MOD) to jointly reconstruct entire scenes, achieving better registration accuracy and reduced inter-object penetration compared to prior work. The dataset, benchmark, code, and pretrained models will be publicly released on the project website.

By Junaid Ahmed Ansari, Ran Ding, Fabio Pizzati, Ivan Laptev
Hugging Face Trending Papers
Jun 22

Humanoid-OmniOcc: Stereo-Based Full-View Occupancy Dataset for Embodied AI

Occupancy prediction at voxel-level granularity is essential for safe robotic navigation and interaction in complex environments. Existing occupancy datasets, however, are predominantly designed for autonomous driving with vehicle-centric biases -- forward-facing cameras, far-field geometry, and static road priors -- limiting their applicability to embodied humanoid perception.