The paper introduces a dual‑backbone architecture for glass segmentation that combines a frozen foundation model with a learned backbone trained on glass‑specific data. By fusing hierarchical multi‑scale features from both backbones, the method produces accurate segmentation masks and achieves state‑of‑the‑art performance on four benchmark datasets. Ablation studies confirm the benefits of the dual‑backbone design and its generalizability across different backbone choices, while also offering competitive inference speeds, especially with lighter backbones.
By Risto Ojala, Tristan Ellison, Mo Chen
Standard depth sensors systematically fail on transparent surfaces, creating corrupted 3D maps and severe navigation hazards. While specialized hardware sensors can detect glass, they lack modularity and have extensive hardware dependencies.
arXiv:2609.36844v1 Announce Type: new
Abstract: Transparent surfaces are ubiquitous in built environments, yet they remain a persistent failure case for robotic perception. RGB cameras perceive the b...
By Suhani Grover, Astik Srivastava, Viswas Dinesh, Avinash Sharma, K. Madhava Krishna
arXiv:2608.13147v2 Announce Type: replace
Abstract: Camera-based autonomous driving perception requires a shared representation that preserves metric 3D structure across synchronized multi-camera str...
By Longfei Xu, Xiaohui Wang, Zehao Huang, Han Li, Ya Yang, Naiyan Wang, Si Liu
CrossDepth introduces geometry-constrained attention for multi-view surround depth estimation, addressing cross-image inconsistencies caused by varying camera intrinsics and limited receptive fields. The method conditions features on per-pixel camera-aware ray embeddings and extends pixel context via cross-image attention limited to geometrically plausible regions. Trained self-supervised with photometric consistency, it achieves better depth accuracy and consistency on DDAD and nuScenes compared to existing self-supervised approaches.
By Samer Abualhanud, Max Mehltretter
PePESeg3D introduces perception priors into a multi‑scale 3D Gaussian segmentation pipeline, integrating monocular depth and mask constraints during geometry reconstruction and dense depth‑color cues with view‑consistent centroid supervision during contrastive feature learning. This dual‑stage approach aligns geometry with semantic structure and compensates for incomplete mask supervision from 2D foundation models. Experiments on SPIn‑NeRF, LERF‑Mask, and NVOS benchmarks show state‑of‑the‑art performance in both multi‑scale segmentation and scene reconstruction.
By Sungjae Choi, Seunghee Koh, Junmo Kim