arXiv Computer Vision
Aug 26

O3N: Omnidirectional Open-Vocabulary Occupancy Prediction for Embodied Intelligent Robotics

O3N is a novel framework that performs open‑vocabulary occupancy prediction from a single omnidirectional RGB image. It introduces a polar‑spiral voxel embedding (PsM) for continuous 360° spatial representation, an Occupancy Cost Aggregation (OCA) module that unifies geometric and semantic supervision, and a Natural Modality Alignment (NMA) pathway that aligns visual, voxel, and text features. Experiments show state‑of‑the‑art results on QuadOcc and Human360Occ benchmarks, with strong cross‑scene generalization and semantic scalability.

By Mengfei Duan, Hao Shi, Fei Teng, Guoqiang Zhao, Yuheng Zhang, Zhiyong Li, Kailun Yang
arXiv Computer Vision
1d ago

VoxelFix: Post-Hoc Semantic Correction of Completed 3D Voxel Maps

VoxelFix is a graph‑based post‑hoc semantic correction method that refines voxel labels in completed 3D voxel maps while preserving their geometry and occupancy. It learns to correct errors by exploiting local geometry and neighboring semantic information, using training pairs generated by corrupting annotated maps with class confusions from upstream perception pipelines. Experiments on OccuFly maps show consistent improvements of 4.23–5.00 percentage points in mIoU, especially for tree, roof, and wall classes, and the method generalizes to out‑of‑distribution aerial scenes.

By Sunesh Praveen Raja Sundarasami, Taehyoung Kim, Johannes Scherer, Toma\v{z} Coti\v{c}, Sivasubiramaniam Subbiah, Andreas Greiner, Paul Spannaus, Sebastian Houben