arXiv:2511.16949v2 Announce Type: replace-cross
Abstract: Dense 3D semantic occupancy perception is critical for mobile robots operating in pedestrian-rich environments, yet it remains underexplored...
By Junseo Kim, Guido Dumont, Xinyu Gao, Gang Chen, Holger Caesar, Javier Alonso-Mora
arXiv:2605.25059v4 Announce Type: replace
Abstract: Crucial for autonomous exploration, online 3D occupancy prediction and mapping incrementally construct dense spatial representations on the fly. Em...
By Ruoyu Wang, Yong Liu, Jiahan Li, Sheng Tao, Yuhang Lin, Yukai Ma
OccAnyScene introduces a unified approach for 3D occupancy prediction across both indoor and outdoor scenes, addressing the challenge of handling diverse camera setups, spatial ranges, voxel specifications, and semantic taxonomies. The method builds on a pretrained depth model, using pixel-aligned frustum feature aggregation and frustum-parameterized Gaussian construction to generate scene-adaptive occupancy predictions. It achieves state-of-the-art performance, scoring 59.92% mIoU on Occ-ScanNet and 23.06% mIoU on SurroundOcc-nuScenes.
By Junjie Liu, Wanshui Gan, Zitong Dai, Guiping Cao, Yan Li, Ke Chen, Dongmei Jiang, Jianguo Zhang, Xiangyuan Lan
arXiv:2603.12789v3 Announce Type: replace
Abstract: Recent advances in 3D foundation models have enabled joint reconstruction of humans and their surrounding environments. However, combining independ...
By Sangmin Kim, Minhyuk Hwang, Geonho Cha, Dongyoon Wee, Jaesik Park
MessyKitchens introduces a new dataset of cluttered real-world kitchen scenes with detailed 3D object shapes, poses, and accurate contact information. The authors extend the SAM 3D single-object reconstruction method with a Multi-Object Decoder (MOD) to jointly reconstruct entire scenes, achieving better registration accuracy and reduced inter-object penetration compared to prior work. The dataset, benchmark, code, and pretrained models will be publicly released on the project website.
By Junaid Ahmed Ansari, Ran Ding, Fabio Pizzati, Ivan Laptev
Occupancy prediction at voxel-level granularity is essential for safe robotic navigation and interaction in complex environments. Existing occupancy datasets, however, are predominantly designed for autonomous driving with vehicle-centric biases -- forward-facing cameras, far-field geometry, and static road priors -- limiting their applicability to embodied humanoid perception.
Embodied tasks demand accurate, flexible, and semantically rich 3D scene representations. 3D semantic occupancy is well suited to this requirement, as it can model holistic 3D spaces by encoding geome...
arXiv:2609.38864v1 Announce Type: new
Abstract: Embodied tasks demand accurate, flexible, and semantically rich 3D scene representations. 3D semantic occupancy is well suited to this requirement, as...
By Jinglong Wang, Yunjie Wang, Zhiyang Zhang, Jiawei He, Ye Yuan, Bo Qiu, Jing Zhang
arXiv:2606. 28215v1 Announce Type: cross Abstract: Extracting dynamic 4D object interactions from massive, in-the-wild monocular videos offers a highly efficient data collection pathway for scaling Embodied AI and training VLAs.
By Jiaxin Li, Yuxiang Wu, Zhenkai Zhang, Xinrui Shi, Haoyuan Wang, Yichen Zhao, Su Linxiang, Chenyang Yu, Mingyu Zhang, Yifan Ding, Boran Wen, Li Zhang, Ruiyang Liu, Yong-Lu Li
arXiv:2609.25654v1 Announce Type: cross
Abstract: Robots operating safely in cluttered everyday environments often need to infer scene geometry from partial observations. Methods that detect objects...
By Dongwon Son, Junhyek Han, Yoontae Cho, Minseok Lee, Hong-seok Choi, Jiwook Choi, Hyungjin Kim, Beomjoon Kim
arXiv:2510. 11014v2 Announce Type: replace-cross Abstract: Autonomous robots often view rooms only partially, through a doorway, where the walls and scene structure hide the geometry and task-relevant semantics needed for safe navigation and goal-directed action.
By Subhransu S. Bhattacharjee, Hao Lu, Dylan Campbell, Rahul Shome
arXiv:2607. 09503v1 Announce Type: cross Abstract: A fundamental challenge in 3D reconstruction and robotic localization is co-visibility: determining which image pairs share overlapping visible surfaces, particularly in scenarios with minimal overlap.
By Filippo Ziliotto, Luciano Serafini, Lamberto Ballan, Tommaso Campari