arXiv Computer Vision By Shiwei Ren, Zhiang Liu, Yongchun Fang, Hongwei Chen

AESplat: Advancing Pose-Free Feed-Forward 3D Gaussian Splatting via Decoupled Appearance Modeling

Read the original on arXiv Computer Vision →

AESplat is a new pose‑free feed‑forward 3D Gaussian Splatting framework that improves rendering quality by decoupling view‑independent and view‑dependent appearance modeling. It directly extracts the base view‑independent appearance from input images and predicts higher‑order spherical harmonic coefficients with a shallow MLP that incorporates 3D‑aware inductive biases. Experiments on several datasets show AESplat outperforms state‑of‑the‑art methods, achieving up to 0.8 dB higher PSNR than NAS3R and 1.1 dB over DepthSplat on RealEstate10K.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Computer Vision
Sep 7

Compact Neural Appearance Models for Efficient Gaussian Splatting

The paper introduces a compact neural appearance model for 3D Gaussian Splatting that replaces traditional low‑order spherical harmonics (SH) with a tiny shared MLP decoding per‑primitive latent codes. It compares SH with recent spherical appearance models, integrating all into a unified CUDA rasterizer and WebGL viewer, and demonstrates that the new neural representation reduces per‑primitive appearance storage from 192 to 28 bytes, speeds optimization by 1.3×, and improves reconstruction quality. The study also analyzes how different appearance parametrizations affect geometry recovery and the handling of non‑static scene content.

By Florian Hahlbohm, Jorge Condor, Linus Franke, Martin Eisemann, Marcus Magnor
arXiv Computer Vision
Sep 21

VoxelTTO: Voxel-Aligned Feed-Forward 3D Gaussian Splatting with Test-Time Optimization

VoxelTTO is a feed‑forward framework that reconstructs 3D Gaussian splatting scenes from multiple images by aggregating dense image features into a global voxel representation and decoding Gaussians from voxel features, thereby eliminating the pixel‑to‑Gaussian correspondence. It incorporates test‑time optimization with lightweight LoRA modules to adapt to known camera parameters while keeping the pretrained visual foundation model frozen. The method replaces standard rasterization with stochastic solid volume rendering, improving geometric fidelity, and demonstrates superior RGB‑D novel‑view synthesis and camera‑pose estimation on Replica, Tanks and Temples, and DTU datasets.

By Yibin Zhao, Yihan Pan, Yangwen Li, Jun Nan, Jianjun Yi
arXiv Computer Vision
Aug 31

WilLaGS: Latent-Conditional 3D Appearance Fields for Robust Gaussian Splatting In-the-Wild

WilLaGS introduces a unified framework that enhances 3D Gaussian Splatting for in-the-wild scenes by learning a continuous global appearance manifold with a β‑VAE and generating dynamic Tri‑Plane features for spatially‑varying local illumination. It also incorporates a self‑supervised perceptual masking mechanism using a Teacher‑Student EMA architecture to suppress transient artifacts and identify inconsistent regions. Experiments on multiple datasets show that WilLaGS achieves state‑of‑the‑art reconstruction quality and novel view synthesis while preserving real‑time rendering efficiency.

By Yuhao Bai, Qianqiu Tan, Lilong Chen, Huanhuan Lv, Lijun Chen