arXiv Computer Vision

AESplat: Advancing Pose-Free Feed-Forward 3D Gaussian Splatting via Decoupled Appearance Modeling

AESplat is a new pose‑free feed‑forward 3D Gaussian Splatting framework that improves rendering quality by decoupling view‑independent and view‑dependent appearance modeling. It directly extracts the base view‑independent appearance from input images and predicts higher‑order spherical harmonic coefficients with a shallow MLP that incorporates 3D‑aware inductive biases. Experiments on several datasets show AESplat outperforms state‑of‑the‑art methods, achieving up to 0.8 dB higher PSNR than NAS3R and 1.1 dB over DepthSplat on RealEstate10K.

arXiv Computer Vision
Sep 7

Compact Neural Appearance Models for Efficient Gaussian Splatting

The paper introduces a compact neural appearance model for 3D Gaussian Splatting that replaces traditional low‑order spherical harmonics (SH) with a tiny shared MLP decoding per‑primitive latent codes. It compares SH with recent spherical appearance models, integrating all into a unified CUDA rasterizer and WebGL viewer, and demonstrates that the new neural representation reduces per‑primitive appearance storage from 192 to 28 bytes, speeds optimization by 1.3×, and improves reconstruction quality. The study also analyzes how different appearance parametrizations affect geometry recovery and the handling of non‑static scene content.

By Florian Hahlbohm, Jorge Condor, Linus Franke, Martin Eisemann, Marcus Magnor
arXiv Computer Vision
Sep 21

VoxelTTO: Voxel-Aligned Feed-Forward 3D Gaussian Splatting with Test-Time Optimization

VoxelTTO is a feed‑forward framework that reconstructs 3D Gaussian splatting scenes from multiple images by aggregating dense image features into a global voxel representation and decoding Gaussians from voxel features, thereby eliminating the pixel‑to‑Gaussian correspondence. It incorporates test‑time optimization with lightweight LoRA modules to adapt to known camera parameters while keeping the pretrained visual foundation model frozen. The method replaces standard rasterization with stochastic solid volume rendering, improving geometric fidelity, and demonstrates superior RGB‑D novel‑view synthesis and camera‑pose estimation on Replica, Tanks and Temples, and DTU datasets.

By Yibin Zhao, Yihan Pan, Yangwen Li, Jun Nan, Jianjun Yi
arXiv Computer Vision
Aug 31

WilLaGS: Latent-Conditional 3D Appearance Fields for Robust Gaussian Splatting In-the-Wild

WilLaGS introduces a unified framework that enhances 3D Gaussian Splatting for in-the-wild scenes by learning a continuous global appearance manifold with a β‑VAE and generating dynamic Tri‑Plane features for spatially‑varying local illumination. It also incorporates a self‑supervised perceptual masking mechanism using a Teacher‑Student EMA architecture to suppress transient artifacts and identify inconsistent regions. Experiments on multiple datasets show that WilLaGS achieves state‑of‑the‑art reconstruction quality and novel view synthesis while preserving real‑time rendering efficiency.

By Yuhao Bai, Qianqiu Tan, Lilong Chen, Huanhuan Lv, Lijun Chen
Hugging Face Trending Papers
Jul 20

FF-ProCams: Feed-Forward Gaussian Splatting for Projector-Camera System

Projector-camera (ProCams) systems achieve active scene perception and controllable appearance manipulation via structured illumination, serving as a core infrastructure for spatial augmented reality, projection mapping, and surface reflectance acquisition. Existing inverse-rendering methods for ProCams deliver high-fidelity results but rely on time-consuming per-scene optimization, while mainstream feed-forward 3D reconstruction models produce baked appearance that cannot adapt to spatially varying projector illumination.

arXiv AI
Sep 10

Aes3D: Aesthetic Assessment in 3D Gaussian Splatting

arXiv:2605.05155v4 Announce Type: replace-cross Abstract: As 3D Gaussian Splatting (3DGS) gains attention in immersive media and digital content creation, assessing the aesthetics of 3D scenes become...

By Chuanzhi Xu, Boyu Wei, Haoxian Zhou, Xuanhua Yin, Zihan Deng, Haodong Chen, Qiang Qu, Weidong Cai
arXiv AI
4d ago

Spackle: Completing Large View Single Image NVS with Adaptive Gaussians

Spackle is a lightweight residual learning framework designed to improve large-view single-image novel view synthesis (NVS) by mitigating capacity competition in hybrid decoupled systems that combine 3D Gaussian Splatting (3DGS) and diffusion models. It operates in three stages: predicting base 3DGS attributes, automatically identifying poorly reconstructed regions, and learning a residual 3DGS focused on those areas. During inference, Spackle merges the baseline and augmented Gaussians to produce high-fidelity novel views, achieving state‑of‑the‑art performance on large-view-deviation cases.

By Xuanzhi Liu, Yuhe Zhou, Xinyi Wu, Zhenyao Wu, Jinghao Chen, Ruize Han, Song Wang