arXiv Computer Vision

Glass Segmentation with Fusion of Learned and General Visual Features

The paper introduces a dual‑backbone architecture for glass segmentation that combines a frozen foundation model with a learned backbone trained on glass‑specific data. By fusing hierarchical multi‑scale features from both backbones, the method produces accurate segmentation masks and achieves state‑of‑the‑art performance on four benchmark datasets. Ablation studies confirm the benefits of the dual‑backbone design and its generalizability across different backbone choices, while also offering competitive inference speeds, especially with lighter backbones.

arXiv Computer Vision
Aug 28

Glass Surface Detection Grounded in 3D Visual Geometry

Glass Surface Detection Grounded in 3D Visual Geometry proposes a new approach that grounds glass surface detection in 3D visual geometry rather than relying solely on 2D appearance cues. The method uses a visual geometry grounded transformer (VGGT) to distill 3D priors and creates glass-aware 3D representations, then applies a multi-task learning framework with a Frequency Self-Attention Module (FSAM) and a Geometry Grounding Block (GeGB) to localize and segment glass surfaces. Experiments show state‑of‑the‑art performance on seven benchmarks, good generalization to video and multi‑modal data, and significant improvements in reconstruction of glass scenes.

By Yiwei Lu, Ke Xu, Tao Yan, Xiaojun Chang, Radu Timofte, Rynson W. H. Lau
arXiv Computer Vision
Sep 25

PePESeg3D: Perception Prior Enhances Multi-Scale Segmentation for 3D Gaussian Splatting

PePESeg3D introduces perception priors into a multi‑scale 3D Gaussian segmentation pipeline, integrating monocular depth and mask constraints during geometry reconstruction and dense depth‑color cues with view‑consistent centroid supervision during contrastive feature learning. This dual‑stage approach aligns geometry with semantic structure and compensates for incomplete mask supervision from 2D foundation models. Experiments on SPIn‑NeRF, LERF‑Mask, and NVOS benchmarks show state‑of‑the‑art performance in both multi‑scale segmentation and scene reconstruction.

By Sungjae Choi, Seunghee Koh, Junmo Kim
arXiv Computer Vision
Sep 18

SenseFuse: Label-Free Fusion of Image and Shape Encoders for Open-Vocabulary 3D Instance Segmentation

SenseFuse introduces a label‑free fusion approach that balances 2D image and 3D shape encoders for open‑vocabulary 3D instance segmentation. By selecting a scene‑level fusion weight through an adaptive, sensitivity‑based mechanism, it improves mask labeling accuracy across multiple datasets, recovering up to 93% of the potential gain from an oracle weight. The method demonstrates that image and shape encoders have complementary failure patterns, leading to higher instance AP in most evaluated settings.

By Euiseok Han, Tri Ton, Hwanhee Kim, Seungyeon Ryu, Chang D. Yoo