arXiv Computer Vision

Glass Surface Detection Grounded in 3D Visual Geometry

Glass Surface Detection Grounded in 3D Visual Geometry proposes a new approach that grounds glass surface detection in 3D visual geometry rather than relying solely on 2D appearance cues. The method uses a visual geometry grounded transformer (VGGT) to distill 3D priors and creates glass-aware 3D representations, then applies a multi-task learning framework with a Frequency Self-Attention Module (FSAM) and a Geometry Grounding Block (GeGB) to localize and segment glass surfaces. Experiments show state‑of‑the‑art performance on seven benchmarks, good generalization to video and multi‑modal data, and significant improvements in reconstruction of glass scenes.

arXiv Computer Vision
Sep 3

Glass Segmentation with Fusion of Learned and General Visual Features

The paper introduces a dual‑backbone architecture for glass segmentation that combines a frozen foundation model with a learned backbone trained on glass‑specific data. By fusing hierarchical multi‑scale features from both backbones, the method produces accurate segmentation masks and achieves state‑of‑the‑art performance on four benchmark datasets. Ablation studies confirm the benefits of the dual‑backbone design and its generalizability across different backbone choices, while also offering competitive inference speeds, especially with lighter backbones.

By Risto Ojala, Tristan Ellison, Mo Chen
arXiv Computer Vision
Sep 7

CrossDepth: Geometry-Constrained Attention for Generalizable Multi-View Surround Depth Estimation

CrossDepth introduces geometry-constrained attention for multi-view surround depth estimation, addressing cross-image inconsistencies caused by varying camera intrinsics and limited receptive fields. The method conditions features on per-pixel camera-aware ray embeddings and extends pixel context via cross-image attention limited to geometrically plausible regions. Trained self-supervised with photometric consistency, it achieves better depth accuracy and consistency on DDAD and nuScenes compared to existing self-supervised approaches.

By Samer Abualhanud, Max Mehltretter
arXiv Computer Vision
Sep 25

PePESeg3D: Perception Prior Enhances Multi-Scale Segmentation for 3D Gaussian Splatting

PePESeg3D introduces perception priors into a multi‑scale 3D Gaussian segmentation pipeline, integrating monocular depth and mask constraints during geometry reconstruction and dense depth‑color cues with view‑consistent centroid supervision during contrastive feature learning. This dual‑stage approach aligns geometry with semantic structure and compensates for incomplete mask supervision from 2D foundation models. Experiments on SPIn‑NeRF, LERF‑Mask, and NVOS benchmarks show state‑of‑the‑art performance in both multi‑scale segmentation and scene reconstruction.

By Sungjae Choi, Seunghee Koh, Junmo Kim
arXiv AI
Jul 24

3D-Aware VLMs with Implicit and Explicit Geometries

arXiv:2607. 21595v1 Announce Type: cross Abstract: Despite rapid progress, most existing vision-language models (VLMs) built from 2D visual inputs often struggle when handling various 3D tasks that require fine-grained spatial understanding and reasoning.

By Wenhao Li, Xueying Jiang, Quanhao Qian, Deli Zhao, Ran Xu, Shijian Lu, Gongjie Zhang
arXiv Computer Vision
Aug 27

3DGS-HPC: Distractor-free 3D Gaussian Splatting with Hybrid Patch-wise Classification

3DGS-HPC is a framework that improves 3D Gaussian Splatting for novel view synthesis by mitigating transient distractors such as moving objects and varying shadows. It combines a patch‑wise classification strategy that uses local spatial consistency for robust region‑level decisions with a hybrid classification metric that adaptively integrates photometric and perceptual cues. Experiments show that this approach outperforms existing methods in reducing distractor effects and enhancing 3DGS quality.

By Jiahao Chen, Yipeng Qin, Ganlong Zhao, Xin Li, Wenping Wang, Guanbin Li