arXiv Computer Vision

LetOccVote: Learning Weakly Supervised 3D Occupancy through Consensus

arXiv Computer Vision
1d ago

VoxelFix: Post-Hoc Semantic Correction of Completed 3D Voxel Maps

VoxelFix is a graph‑based post‑hoc semantic correction method that refines voxel labels in completed 3D voxel maps while preserving their geometry and occupancy. It learns to correct errors by exploiting local geometry and neighboring semantic information, using training pairs generated by corrupting annotated maps with class confusions from upstream perception pipelines. Experiments on OccuFly maps show consistent improvements of 4.23–5.00 percentage points in mIoU, especially for tree, roof, and wall classes, and the method generalizes to out‑of‑distribution aerial scenes.

By Sunesh Praveen Raja Sundarasami, Taehyoung Kim, Johannes Scherer, Toma\v{z} Coti\v{c}, Sivasubiramaniam Subbiah, Andreas Greiner, Paul Spannaus, Sebastian Houben
arXiv Computer Vision
1d ago

Object Concepts Emerge from Motion

arXiv:2609.04348v1 Announce Type: new Abstract: Object-centric visual representations are important for physical-world perception, but existing visual pretraining methods often capture semantic categ...

By Boshi Li, Xiaohui Wang, Xiaoyang Wu, Zhichao Li, Ya Yang, Naiyan Wang
Hugging Face Trending Papers
Jul 30

MonoVoc: Decoupling Geometry and Semantics for Lightweight Monocular Open-Vocabulary 3D Gaussians

Open vocabulary 3D scene understanding is essential for next-generation interactive systems, empowering users to intuitively query and navigate reconstructed environments using natural language. However, current 3D Gaussian frameworks are often bottlenecked by restrictive multiview capture requirements, costly scene-specific optimization, and the massive memory overhead of storing dense language features.