ExtrinSplat is a new framework that separates geometry from semantics in 3D Gaussian Splatting scenes. It clusters Gaussians into overlapping 3D object groups and uses a Vision‑Language Model to generate lightweight textual hypotheses, creating an extrinsic index layer that handles complex polysemy. This approach reduces adaptation time from hours to minutes, cuts storage overhead by orders of magnitude, and outperforms existing embedding‑based methods on open‑vocabulary 3D object selection and semantic segmentation benchmarks.
By Jiayu Ding, Xinpeng Liu, Zhiyi Pan, Shiqiang Long, Ge Li
GaussianDS introduces a depth‑supervised framework for 3D Gaussian Splatting that jointly optimizes RGB appearance, depth, and compact semantics from scratch. By arranging multi‑view images into a pose‑aware pseudo‑video and propagating view‑consistent masks via SAM2, the method aligns semantic lifting with geometric cues, using depth supervision and edge‑aware refinement to curb semantic drift and boundary leakage. The approach achieves state‑of‑the‑art performance on LERF and 3D‑OVS benchmarks while preserving high‑fidelity reconstruction and enabling downstream tasks such as 3D object removal.
By Yufei Zhang, Chenlu Zhan, Hongwei Wang
arXiv:2609.15137v1 Announce Type: cross
Abstract: 3D Gaussian language fields provide an explicit, spatially grounded representation for 3D visual question answering (VQA), but their dense semantic f...
By Davit Soselia, Joseph JaJa, Amitabh Varshney
Feed-forward 3D Gaussian Splatting (3DGS) enables scalable scene reconstruction without per-scene optimization, yet produces dense Gaussians that are costly to store and transmit. Existing feed-forward Gaussian compression methods formulate decoding as deterministic representation recovery, which becomes inadequate at low bitrates when high-frequency textures and view-dependent appearance are discarded.
The paper introduces WINGS, a reference‑free Gaussian splatting inpainting technique that operates directly in 3D. It uses a large, pre‑trained 3D generative prior and a structure completion network to reconstruct missing geometry and appearance without generating reference views. The method claims faster performance and reduced multi‑view inconsistency compared to 2D diffusion‑based approaches, and its effectiveness is validated through experiments and a user study.
By No\'e Lallouet, Michael Fischer, Elie Michel
Open vocabulary 3D scene understanding is essential for next-generation interactive systems, empowering users to intuitively query and navigate reconstructed environments using natural language. However, current 3D Gaussian frameworks are often bottlenecked by restrictive multiview capture requirements, costly scene-specific optimization, and the massive memory overhead of storing dense language features.
PePESeg3D introduces perception priors into a multi‑scale 3D Gaussian segmentation pipeline, integrating monocular depth and mask constraints during geometry reconstruction and dense depth‑color cues with view‑consistent centroid supervision during contrastive feature learning. This dual‑stage approach aligns geometry with semantic structure and compensates for incomplete mask supervision from 2D foundation models. Experiments on SPIn‑NeRF, LERF‑Mask, and NVOS benchmarks show state‑of‑the‑art performance in both multi‑scale segmentation and scene reconstruction.
By Sungjae Choi, Seunghee Koh, Junmo Kim
The paper introduces a fusion‑aware hierarchical Gaussian patch representation that enables direct class‑guided generation of 3D Gaussian Splatting (3DGS) objects. By decomposing irregular Gaussian sets into canonical local patches and encoding them as structured tokens, the method fuses global class semantics with patch‑level geometry, appearance, spatial correspondence, and rendering‑sensitive cues. A structure‑aware rectified flow model, conditioned on patch positions and coupled with global‑local velocity prediction and density‑aware weighting, produces class‑conditioned 3DGS objects within seconds, achieving more coherent geometry, sharper local details, and better multi‑view consistency than baseline models.
By Yizhao Wang, Jingbo Wang, Guantao Zhang
arXiv:2609.12343v1 Announce Type: new
Abstract: Feed-forward Gaussian splatting models have demonstrated remarkable effectiveness in reconstructing three-dimensional (3D) objects from a few two-dimen...
By Yunsu Jeong, Hyuk Heo, Youngsang Kwak, Jaehwa Kwak, Il Yong Chun
arXiv:2607. 00746v1 Announce Type: cross Abstract: The bird's-eye view (BEV) representation enables multi-sensor features to be fused within a unified space, serving as the primary approach for achieving comprehensive 3D perception.
By Xiao Zhao, Chang Liu, Mingxu Zhu, Zheyuan Zhang, Linna Song, Qingliang Luo, Chufan Guo, Kuifeng Su
Spackle is a lightweight residual learning framework designed to improve large-view single-image novel view synthesis (NVS) by mitigating capacity competition in hybrid decoupled systems that combine 3D Gaussian Splatting (3DGS) and diffusion models. It operates in three stages: predicting base 3DGS attributes, automatically identifying poorly reconstructed regions, and learning a residual 3DGS focused on those areas. During inference, Spackle merges the baseline and augmented Gaussians to produce high-fidelity novel views, achieving state‑of‑the‑art performance on large-view-deviation cases.
By Xuanzhi Liu, Yuhe Zhou, Xinyi Wu, Zhenyao Wu, Jinghao Chen, Ruize Han, Song Wang
Recent advancements in 3D Gaussian Splatting (3DGS) have enabled language-guided scene understanding. However, existing Referring 3D Gaussian Splatting (R3DGS) methods are fundamentally restricted to single-target queries.