arXiv Computer Vision

What Builds the Scene? Luminance Dominates Geometry Formation in 3D Gaussian Splatting

arXiv Machine Learning
Jul 1

Drop-In Perceptual Optimization for 3D Gaussian Splatting

arXiv:2603. 23297v2 Announce Type: replace-cross Abstract: Despite their output being ultimately consumed by human viewers, 3D Gaussian Splatting (3DGS) methods often rely on ad-hoc combinations of pixel-level losses, resulting in blurry renderings.

By Ezgi Ozyilkan, Zhiqi Chen, Oren Rippel, Jona Ball\'e, Kedar Tatwawadi
Hugging Face Trending Papers
Jul 23

Engine-Native Editable 3D World Reconstruction with Objects and Lighting

Editable 3D scene creation requires object instances and lights that can be inspected, moved, and imported into standard engines, yet existing single-image methods largely stop at room-scale geometry, baked/global illumination, or text-driven generation. We introduce Lumera (Light-aware Unified Engine-native Reconstruction and Assembly), a benchmark and reference pipeline for engine-native, light-aware 3D scene parsing from a single image.

arXiv Computer Vision
6d ago

Gauss What You Need: Compact Gaussian Splatting Across Scene Scales

Gauss What You Need: Compact Gaussian Splatting Across Scene Scales introduces TangoGS, a method that automatically selects the number of Gaussian primitives for 3D Gaussian Splatting by combining capture-derived model sizing with training-based adaptation. The approach first estimates a learning allowance based on the capture’s total pixels, then adjusts the number of Gaussians during training according to reconstruction quality. On standard benchmarks, TangoGS matches the best baseline’s PSNR while using 48% fewer Gaussians, and on larger captures it scales automatically to achieve the highest mean PSNR with 2.3× more Gaussians.

By Afif Boudaoud, Jiayi Liu, Alexandru Calotoiu, Torsten Hoefler
arXiv Computer Vision
Sep 7

Compact Neural Appearance Models for Efficient Gaussian Splatting

The paper introduces a compact neural appearance model for 3D Gaussian Splatting that replaces traditional low‑order spherical harmonics (SH) with a tiny shared MLP decoding per‑primitive latent codes. It compares SH with recent spherical appearance models, integrating all into a unified CUDA rasterizer and WebGL viewer, and demonstrates that the new neural representation reduces per‑primitive appearance storage from 192 to 28 bytes, speeds optimization by 1.3×, and improves reconstruction quality. The study also analyzes how different appearance parametrizations affect geometry recovery and the handling of non‑static scene content.

By Florian Hahlbohm, Jorge Condor, Linus Franke, Martin Eisemann, Marcus Magnor
arXiv Computer Vision
Sep 25

Only What Was Seen: Observation-Gram Compaction of View-Dependent Appearance in 3D Gaussian Splatting

The paper introduces an observation‑Gram matrix that captures how each Gaussian in a 3D Gaussian Splatting model is viewed from training camera directions. This matrix serves as a distortion metric, enabling closed‑form degree reduction, Lagrangian rate‑distortion degree allocation, and matrix‑weighted vector quantisation. When applied to the Compressed3D framework, the metric improves PSNR by 0.49 dB before fine‑tuning and still yields a 0.32 dB gain at matched bitrate without any training images, while a training‑free stack built on the metric is 15% smaller than the image‑free GSICO at equal quality on Mip‑NeRF 360.

By Krzysztof Pietroszek
arXiv AI
Sep 7

Where Appearance Fails, Geometry Recognizes: A CAD-Free 3D Shape Prior That Complements Vision Foundation Models

The paper introduces a CAD‑free 3D shape prior that enhances object recognition by reconstructing each object with 3D Gaussian Splatting (3DGS) from short RGB‑D scans and fusing the resulting shape prototype with frozen DINOv2 image features. Experiments on T‑LESS and HOPE datasets show that geometry alone can match or exceed CAD‑based recognition, and that the combined approach improves performance, especially on shape‑distinctive or partially occluded objects. The study demonstrates that the benefit comes from the geometric information rather than rendered pixels, and that the prior is complementary to frozen vision features.

By Chenxi Tao, Seung-Kyum Choi