arXiv Computer Vision

RoomLight: A 2.5D Illumination Prior for Indoor Environments

arXiv Computer Vision
Sep 24

GLOW: Global Illumination-Aware Inverse Rendering of Indoor Scenes Captured with Dynamic Co-Located Light & Camera

GLOW is a Global Illumination‑aware inverse rendering framework for indoor scenes captured with dynamic co‑located light and camera setups. It combines a neural implicit surface representation with a neural radiance cache to jointly optimize geometry and reflectance, while introducing a dynamic radiance cache and a surface‑angle‑weighted radiometric loss to handle near‑field motion, strong inter‑reflections, and specular highlights. Experiments demonstrate that GLOW significantly outperforms prior methods in estimating material reflectance under both natural and co‑located illumination.

By Jiaye Wu, Saeed Hadadan, Geng Lin, Peihan Tu, Auguste Gezalyan, Matthias Zwicker, David Jacobs, Roni Sengupta
arXiv Computer Vision
Sep 11

Shedding Light: A Benchmark for Evaluating Lighting Understanding in Generative Image Models

The paper introduces a benchmark called Shedding Light to evaluate how well generative image models understand and reproduce lighting. The benchmark tests models by asking them to inpaint a simple object, called a light probe, into real photographs and then compares the generated probe to the ground truth to assess lighting direction, colour, and radiance. The authors provide a scalable protocol and open-source code and data for systematic assessment of photometric accuracy in future models.

By Justine Giroux, Jack Oliver Hilliard, Yannick Hold-Geoffroy, Javier Vazquez-Corral, Jean-Fran\c{c}ois Lalonde
arXiv AI
Sep 10

RelightFormer: Feed-forward Generative Transformer for Multiview Object Relighting

RelightFormer is a feed‑forward generative Transformer that performs single‑ and multi‑view image relighting without explicit intrinsic property estimation. It incorporates a latent illumination module that injects target environment maps into spatial features via cross‑attention, and uses permutation‑invariant positional encodings to process unordered multi‑view inputs symmetrically. Trained on the large Laval Objaverse Dataset, the model achieves state‑of‑the‑art visual and photorealistic relighting quality, and demonstrates strong zero‑shot generalization across various relighting tasks.

By Hejun Wang, Jinxi Li, Junwei Jiang, Shiwei Mao, Hu Cheng, Shouwang Huang, Bo Yang
arXiv AI
Sep 21

Physically Based Rendering in the Latent Space

The paper introduces a method that applies physically based rendering (PBR) within the latent space of variational autoencoders used in image diffusion models. By modifying the rendering equation and using a differentiable renderer, the authors can generate latent maps that guide content creation with physically accurate lighting. The approach is trained on a single rendered image and then shown to generalize to changes in scene geometry, lighting, and camera viewpoint.

By Vuk Radovanovic, Vishesh Gupta, Adrien Gruson, Binh-Son Hua
arXiv AI
Sep 10

WildRelight: A Real-World Benchmark and Physics-Guided Adaptation for Single-Image Relighting

WildRelight is the first in-the-wild dataset designed to evaluate single-image relighting models, featuring high-resolution outdoor scenes captured under strictly aligned, temporally varying natural illuminations paired with high-dynamic-range environment maps. The benchmark demonstrates that state-of-the-art models trained on synthetic data suffer severe domain shifts when applied to real-world imagery. Leveraging the dataset’s temporal structure, the authors introduce a physics-guided inference framework combining Diffusion Posterior Sampling with Temporal Sampling-Aware Test-Time Adaptation, enabling synthetic models to self-supervise and align with real-world statistics on-the-fly.

By Lezhong Wang, Mehmet Onurcan Kaya, Siavash Bigdeli, Jeppe Revall Frisvad
arXiv Computer Vision
Sep 17

PureLight: Learning Complex Luminaires with Light Tracing

PureLight introduces a neural approach to estimate the appearance of complex luminaires that are difficult for traditional path tracing, such as small emitters surrounded by multiple specular layers. The method uses light tracing to build paths from emitters to exit surfaces and learns the probability density function of outgoing radiance with a large normalizing flow network, then distills this into a lightweight MLP for efficient inference. Additionally, a sampling network and a blending network are trained to compute direct illumination and composite the luminaire into arbitrary scenes, enabling low‑sample rendering of challenging luminaires.

By Pedro Figueiredo, Zixuan Li, Beibei Wang, Milo\v{s} Ha\v{s}an, Nima Khademi Kalantari
arXiv Machine Learning
Sep 4

Sparse auto-regressive modeling for scene generation from multi-view images

The paper introduces SPAR3S, a sparse voxel‑aligned 3D latent generative model that completes 3D scenes from sparse, unconstrained multi‑view images. It learns a compact voxel‑aligned latent space using photometric supervision via differentiable 3D Gaussian Splatting, and employs a masked autoregressive transformer to predict missing voxel occupancy and latent tokens. Experiments on synthetic indoor scenes and RealEstate10k show that SPAR3S achieves higher novel‑view quality than prior methods and generalizes to real‑world data.

By Thomas Lucas, Maxime Pietrantoni, Philippe Weinzaepfel, Wonjune Cho, Bardienus Pieter Duisterhof, Vincent Leroy, Jerome Revaud
Hugging Face Trending Papers
Jun 28

EvLIR: Learning Illumination Residuals from Ordered Events for Low-Light Image Enhancement

Low-light image enhancement is severely ill-posed when the input frame contains missing structure, saturated noise, and weak local contrast. Event cameras provide asynchronous brightness-change observations with high temporal resolution, but prior works often treat voxel channels as an unordered or static feature stack before fusion, rather than explicitly modeling their within-window temporal evolution, weakening the temporal evidence that makes events useful.

Hugging Face Trending Papers
Sep 3

Sparse auto-regressive modeling for scene generation from multi-view images

The paper introduces SPAR3S, a sparse voxel‑aligned 3D latent generative model that completes 3D scenes from sparse, unconstrained multi‑view images. By representing only occupied voxels in a compact latent space and training a masked autoregressive transformer with photometric supervision via differentiable 3D Gaussian Splatting, the method predicts missing latent tokens and spatial support, enabling efficient and spatially consistent generation of unseen regions. Experiments on synthetic indoor scenes and RealEstate10k demonstrate higher novel‑view quality and real‑world applicability compared to prior work.