arXiv Computer Vision By Zhikun Zhou, Kunyu Peng, Runyi Yang, Junhao Cai, Di Wen, Ruiping Liu, Danda Pani Paudel, Yi Zhou, Luc Van Gool, Kailun Yang

CoRef-GS: Cooperative Referring Gaussian Splatting for Multi-Agent Scene Understanding

Read the original on arXiv Computer Vision →

CoRef-GS introduces a cooperative referring Gaussian splatting framework for multi‑agent scene understanding, enabling robots to ground object‑ and relation‑centric language queries across independently reconstructed maps. The method builds local open‑vocabulary instance‑aware Gaussian maps, aligns them using a cross‑agent module that enforces geometric and semantic consistency, and grounds queries with a view‑conditioned mask relation graph. Experiments on a new dual‑quadruped benchmark show significant improvements, reducing rotation error from 2.58° to 0.15° and raising real‑world referring mIoU from 52.6% to 68.8%.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Computer Vision
Sep 16

DGSG-Mind: Dynamic 3D Gaussian Scene Graphs for Long-Term Scene Understanding and Grounding

DGSG-Mind introduces a hybrid instance-aware 3D Gaussian dynamic scene graph system that integrates open‑vocabulary semantic information into dynamic 3D scene representations. By coupling a probabilistic voxel grid with explicit 3D Gaussians, it achieves robust cross‑modal instance fusion, incremental semantic mapping, and dynamic change handling through Gaussian‑based relocalization and masked refinement. The system builds a hierarchical scene graph and a 3D Gaussian Mind for multimodal reasoning, achieving state‑of‑the‑art zero‑shot 3D visual grounding and strong performance in open‑vocabulary semantic segmentation and scene reconstruction, and is demonstrated on real‑world robots.

By Luzhou Ge, Xiangyu Zhu, Jinyan Liu, Xuesong Li
arXiv Computer Vision
Sep 14

AnchorVLN: Geometry-Anchored Vision-Language Grounding Reasoning for Open-Vocabulary Navigation

AnchorVLN is an open‑vocabulary vision‑language navigation system that separates semantic proposals from geometric metrics. It uses a VLM to generate semantics while a geometry module supplies reliable metric quantities such as range and bearing, all within a Model Context Protocol server. The system achieves 64.4% on instruction following and improves object‑reference accuracy, reducing median center error from 3.37 m to 2.48 m.

By Long Giang Vu, Chengkai Yao, Yuxin Liu, FNU Aryan, Rajath Chandrashekar Aralikatti
arXiv AI
Jun 24

G$^3$VLA: Geometric inductive bias for Vision-Language-Action Models

arXiv:2606. 24472v1 Announce Type: cross Abstract: Vision-language-action (VLA) models have made rapid progress in generalist robot manipulation by harnessing semantic knowledge from pretrained vision-language backbones, but their visual tokens remain grounded in 2D image coordinates rather than the calibrated geometry of the robot's cameras -- a mismatch especially pronounced in multi-camera setups, where views are coupled by known intrinsics and extrinsics yet processed as independent images.

By Yue Peng, Yongzhe Zhao, Artur Habuda, Khuyen Pham, Yanheng Zhu, Tran Nguyen Le, Fares Abu-Dakka, Li Guo