arXiv Computer Vision

Privacy-Preserving Semantic Segmentation from High-Resolution Depth and Ultra-Low-Resolution RGB

The paper introduces a privacy‑preserving approach for semantic segmentation that fuses high‑resolution depth with ultra‑low‑resolution RGB images. A joint 2D framework uses depth to guide RGB reconstruction and RGB‑D segmentation, while an end‑to‑end 2D‑to‑3D pipeline consolidates 2D features for 3D segmentation. Experiments on ScanNet demonstrate superior 2D and 3D performance compared to other privacy‑preserving methods, strong zero‑shot transfer to SUN RGB‑D and SceneNN, and reduced recoverability of sensitive data, with real‑robot tests showing effective object‑goal navigation.

arXiv Computer Vision
Sep 18

Depth-Only Open-Vocabulary 3D Semantic Segmentation For Privacy-Preserving Robotic Applications

The paper introduces a privacy‑preserving approach to open‑vocabulary 3D semantic segmentation that operates solely on depth data, eliminating the use of RGB images to avoid disclosing scene‑specific visual information. It proposes a stricter depth‑only evaluation protocol and presents UTTO, a model‑agnostic uncertainty‑guided test‑time optimization framework that refines predictions from frozen open‑vocabulary 3D backbones using structured predictive uncertainty. Experiments on ScanNet and Matterport3D show consistent improvements, and additional analyses demonstrate the method’s relevance for privacy‑constrained robotic applications.

By Xuying Huang, Sicong Pan, Maren Bennewitz
arXiv Computer Vision
Sep 24

GaussianDS: Depth-supervised Semantic Gaussian Splatting for Scene Understanding

GaussianDS introduces a depth‑supervised framework for 3D Gaussian Splatting that jointly optimizes RGB appearance, depth, and compact semantics from scratch. By arranging multi‑view images into a pose‑aware pseudo‑video and propagating view‑consistent masks via SAM2, the method aligns semantic lifting with geometric cues, using depth supervision and edge‑aware refinement to curb semantic drift and boundary leakage. The approach achieves state‑of‑the‑art performance on LERF and 3D‑OVS benchmarks while preserving high‑fidelity reconstruction and enabling downstream tasks such as 3D object removal.

By Yufei Zhang, Chenlu Zhan, Hongwei Wang
arXiv AI
Jun 8

MatterDoor: Sampling Zero-shot Spatio-semantic Priors using Generative Models

arXiv:2510. 11014v2 Announce Type: replace-cross Abstract: Autonomous robots often view rooms only partially, through a doorway, where the walls and scene structure hide the geometry and task-relevant semantics needed for safe navigation and goal-directed action.

By Subhransu S. Bhattacharjee, Hao Lu, Dylan Campbell, Rahul Shome
arXiv Computer Vision
Sep 25

RGBD20K: A Large-Scale Benchmark for RGB-D Semantic Segmentation

RGBD20K is a new large-scale RGB‑D semantic segmentation dataset featuring 20,000 image pairs and 160 fine‑grained categories, surpassing existing benchmarks like NYUv2 and SUN RGB‑D in both scale and semantic diversity. The dataset provides high‑fidelity annotations obtained through rigorous re‑evaluation and correction of prior labels, ensuring a clean ground‑truth foundation. Additionally, the authors introduce a score‑purified fusion (SPF) method that achieves state‑of‑the‑art performance across evaluated benchmarks, demonstrating the value of high‑quality multimodal information.

By Shaohua Dong, Zexuan Meng, Haiyan Sun, Bing Fan, Cuicui Zhang, Dylan Joseph, Kewei Sha, Yunhe Feng, Heng Fan
Hugging Face Trending Papers
Sep 24

RGBD20K: A Large-Scale Benchmark for RGB-D Semantic Segmentation

RGBD20K is a new large-scale RGB‑D dataset designed to advance semantic segmentation research. It contains 20,000 image pairs annotated with 160 fine‑grained categories, far exceeding the diversity of existing benchmarks such as NYUv2 and SUN RGB‑D. The authors also provide high‑fidelity annotations and introduce a score‑purified fusion (SPF) method that achieves state‑of‑the‑art results on multiple benchmarks.

arXiv AI
Jul 1

MVP-Nav: Multi-layer Value Map Planner Navigator

arXiv:2606. 31919v1 Announce Type: cross Abstract: Zero-shot Object Goal Navigation (ZSON) with RGB-only perception poses a fundamental challenge for embodied agents, as the absence of explicit depth information introduces severe physical uncertainty and semantic-physical misalignment.

By Wenyuan Xie, Shaokai Wu, Yijin Zhou, Yanbiao Ji, Guodong Zhang, Bayram Bayramli, Qiuchang Li, Xunchu Zhou, Yue Ding, Hongtao Lu
arXiv AI
4d ago

UniAfford: Token-Routed Multitask Learning for Generalizable 2D-3D Affordance Perception

arXiv:2609.37264v1 Announce Type: cross Abstract: Affordance perception aims to localize actionable regions supporting embodied interaction, yet 2D and 3D affordance grounding have evolved as separat...

By Yuhao Liu, Yiming Zhong, Hanqing Wang, Shaocheng Yan, Yuhang Zhang, Wenzhou Lyu, Ziyang Ding, Wei Zhang, Xue Zhao, Jin Pan, Yuexin Ma, Xinge Zhu
arXiv AI
Jul 7

MetricAnything: Scaling Metric Depth Pretraining with Noisy Heterogeneous Sources

arXiv:2601. 22054v2 Announce Type: replace-cross Abstract: Scaling has powered recent advances in vision foundation models, yet extending this paradigm to metric depth estimation remains challenging due to heterogeneous sensor noise, camera-dependent biases, and metric ambiguity in noisy cross-source 3D data.

By Baorui Ma, Jiahui Yang, Donglin Di, Xuancheng Zhang, Jianxun Cui, Hao Li, Yan Xie, Wei Chen