arXiv AI By Oguzhan Baser, Mirac Sozen, Kaan Kale, Sandeep Chinchali, Sriram Vishwanath

RoboShape: Information-Theoretic Point Cloud Representations for Privacy-Aware Robot Perception

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

arXiv Computer Vision
Sep 18

Depth-Only Open-Vocabulary 3D Semantic Segmentation For Privacy-Preserving Robotic Applications

The paper introduces a privacy‑preserving approach to open‑vocabulary 3D semantic segmentation that operates solely on depth data, eliminating the use of RGB images to avoid disclosing scene‑specific visual information. It proposes a stricter depth‑only evaluation protocol and presents UTTO, a model‑agnostic uncertainty‑guided test‑time optimization framework that refines predictions from frozen open‑vocabulary 3D backbones using structured predictive uncertainty. Experiments on ScanNet and Matterport3D show consistent improvements, and additional analyses demonstrate the method’s relevance for privacy‑constrained robotic applications.

By Xuying Huang, Sicong Pan, Maren Bennewitz
arXiv Computer Vision
Sep 24

Privacy-Preserving Semantic Segmentation from High-Resolution Depth and Ultra-Low-Resolution RGB

The paper introduces a privacy‑preserving approach for semantic segmentation that fuses high‑resolution depth with ultra‑low‑resolution RGB images. A joint 2D framework uses depth to guide RGB reconstruction and RGB‑D segmentation, while an end‑to‑end 2D‑to‑3D pipeline consolidates 2D features for 3D segmentation. Experiments on ScanNet demonstrate superior 2D and 3D performance compared to other privacy‑preserving methods, strong zero‑shot transfer to SUN RGB‑D and SceneNN, and reduced recoverability of sensitive data, with real‑robot tests showing effective object‑goal navigation.

By Xuying Huang, Swithinraj Moses Daniel, Sicong Pan, Sebastian Houben, Maren Bennewitz
arXiv Computer Vision
Aug 26

O3N: Omnidirectional Open-Vocabulary Occupancy Prediction for Embodied Intelligent Robotics

O3N is a novel framework that performs open‑vocabulary occupancy prediction from a single omnidirectional RGB image. It introduces a polar‑spiral voxel embedding (PsM) for continuous 360° spatial representation, an Occupancy Cost Aggregation (OCA) module that unifies geometric and semantic supervision, and a Natural Modality Alignment (NMA) pathway that aligns visual, voxel, and text features. Experiments show state‑of‑the‑art results on QuadOcc and Human360Occ benchmarks, with strong cross‑scene generalization and semantic scalability.

By Mengfei Duan, Hao Shi, Fei Teng, Guoqiang Zhao, Yuheng Zhang, Zhiyong Li, Kailun Yang