arXiv Computer Vision

Dior: Drawing the Light of Image via Material-Decoupled Illumination Representation

Hugging Face Trending Papers
Jun 21

Generative Relightable Avatars

We present Generative Relightable Avatars (GRA), a person-specific method for photorealistic free-view rendering and environment-map relighting of full-body humans. We postulate that modeling fine-grained appearance details is inherently a one-to-many problem that can benefit from a generative formulation.

arXiv AI
Jul 1

Intrinsic decomposition and editing of 3D Gaussian splats

arXiv:2606. 31637v1 Announce Type: cross Abstract: Intrinsic decomposition which expresses image colors as the product of diffuse albedo and shading, possibly augmented with view-dependent residuals has a long history in image editing as it enables the modification of object colors and textures without altering lighting.

By Alexandre Lanvin, Jeffrey Hu, Simon Lucas, Adrien Bousseau, George Drettakis
Hugging Face Trending Papers
Jul 7

From RGB Generation to Dense Field Readout: Pixel-Space Dense Prediction with Text-to-Image Models

Large-scale text-to-image models are attractive backbones for dense prediction because RGB generation pretraining learns rich semantic, structural, and geometric priors. Existing generative and editing approaches reuse these priors by casting dense prediction as target generation: annotations such as depth, normals, alpha mattes, masks, and heatmaps are encoded into an RGB-trained VAE latent space and decoded back as image-like targets.

arXiv AI
Aug 26

Luce: Relightable Gaussians for 3D Asset Generation

Luce is a 3D representation that unifies geometry and physically based rendering (PBR) materials within a voxelized multimodal Gaussian cloud, using dedicated Gaussian primitives for each modality. A variational autoencoder compresses this representation into a unified material‑aware latent space, which a rectified‑flow transformer generates from a single image conditioned on multi‑layer features from a pretrained image encoder. The latent decodes into relightable PBR Gaussians and an optional textured mesh with a tangent‑space normal map, achieving state‑of‑the‑art single‑image‑to‑3D generation on Toys4K and improving CLIP image‑alignment scores on a benchmark of AI‑generated images.

By Mayank Singh, Michele Stoppa, Alvise Memo, Rui Yu, Harsha Kalli, Srimanth Gunturi, Muhammad Ahmed Riaz, Behrooz Shahsavari, Waleed Abdulla, David E. Jacobs