Glass Surface Detection Grounded in 3D Visual Geometry proposes a new approach that grounds glass surface detection in 3D visual geometry rather than relying solely on 2D appearance cues. The method uses a visual geometry grounded transformer (VGGT) to distill 3D priors and creates glass-aware 3D representations, then applies a multi-task learning framework with a Frequency Self-Attention Module (FSAM) and a Geometry Grounding Block (GeGB) to localize and segment glass surfaces. Experiments show state‑of‑the‑art performance on seven benchmarks, good generalization to video and multi‑modal data, and significant improvements in reconstruction of glass scenes.
By Yiwei Lu, Ke Xu, Tao Yan, Xiaojun Chang, Radu Timofte, Rynson W. H. Lau
arXiv:2512.15708v2 Announce Type: replace
Abstract: Foundation models are vital tools in various Computer Vision applications. They take as input a single RGB image and output a deep feature represen...
By Leo Segre, Or Hirschorn, Shai Avidan
arXiv:2607. 05568v1 Announce Type: cross Abstract: Representing 3D shapes as compact sets of geometric primitives is fundamental to robotics, simulation, and scene understanding.
By Gregor Kobsik, Tim Elsner, Leif Kobbelt
PePESeg3D introduces perception priors into a multi‑scale 3D Gaussian segmentation pipeline, integrating monocular depth and mask constraints during geometry reconstruction and dense depth‑color cues with view‑consistent centroid supervision during contrastive feature learning. This dual‑stage approach aligns geometry with semantic structure and compensates for incomplete mask supervision from 2D foundation models. Experiments on SPIn‑NeRF, LERF‑Mask, and NVOS benchmarks show state‑of‑the‑art performance in both multi‑scale segmentation and scene reconstruction.
By Sungjae Choi, Seunghee Koh, Junmo Kim
Standard depth sensors systematically fail on transparent surfaces, creating corrupted 3D maps and severe navigation hazards. While specialized hardware sensors can detect glass, they lack modularity and have extensive hardware dependencies.
arXiv:2609.22896v1 Announce Type: new
Abstract: Autonomous vehicles operating in open-world scenarios are inevitably confronted with previously unknown objects, such as exotic animals or loose cargo....
By Serin Varghese, Fabian H\"uger, Kira Maag
arXiv:2609.24226v1 Announce Type: new
Abstract: Instance segmentation is a fundamental computer vision task with diverse real-world applications. Recently, prompt-driven foundation models have shown...
By Lufei Liu, Guojie Li, Suncheng Xiang, Fan Zhang
SenseFuse introduces a label‑free fusion approach that balances 2D image and 3D shape encoders for open‑vocabulary 3D instance segmentation. By selecting a scene‑level fusion weight through an adaptive, sensitivity‑based mechanism, it improves mask labeling accuracy across multiple datasets, recovering up to 93% of the potential gain from an oracle weight. The method demonstrates that image and shape encoders have complementary failure patterns, leading to higher instance AP in most evaluated settings.
By Euiseok Han, Tri Ton, Hwanhee Kim, Seungyeon Ryu, Chang D. Yoo
arXiv:2609.22687v1 Announce Type: new
Abstract: We present PanoSeg3R, a feed-forward framework for 3D panoramic semantic segmentation. Unlike existing methods designed for perspective inputs, PanoSeg...
By Heechan Yoon, Dongki Jung, Phuc Nguyen, Ming Lin, Dinesh Manocha
arXiv:2609.36844v1 Announce Type: new
Abstract: Transparent surfaces are ubiquitous in built environments, yet they remain a persistent failure case for robotic perception. RGB cameras perceive the b...
By Suhani Grover, Astik Srivastava, Viswas Dinesh, Avinash Sharma, K. Madhava Krishna
arXiv:2608.31052v1 Announce Type: cross
Abstract: Semantic segmentation decomposes an image into distinct mask regions corresponding to different object categories, such as people, cars, signs or bui...
By Keith G. Mills, Evan B. Sanders, Gregory J. Matthews, Juliet K. Brophy
arXiv:2505.15147v3 Announce Type: replace
Abstract: Remote sensing images (RSIs) capture both natural and human-induced changes on the Earth's surface. Semantic segmentation (SS) of RSIs enables the...
By Quanwei Liu, Tao Huang, Jiaqi Yang, Wei Xiang