arXiv:2606. 30380v1 Announce Type: cross Abstract: We present RenderFormer++, a scalable and physically grounded feed-forward neural rendering framework for global illumination in mesh scenes.
By Huangsheng Du, Haoran Zhu, Youcheng Cai, Jinyang Meng, Ligang Liu
arXiv:2609.05738v1 Announce Type: cross
Abstract: We present 'RenderFormer-V2', a unified learned transformer-based neural rendering model, complementary to modern physics-based rendering systems, th...
By Chong Zeng, Yue Dong, Pieter Peers, Lvmin Zhang, Maneesh Agrawala
RelightFormer is a feed‑forward generative Transformer that performs single‑ and multi‑view image relighting without explicit intrinsic property estimation. It incorporates a latent illumination module that injects target environment maps into spatial features via cross‑attention, and uses permutation‑invariant positional encodings to process unordered multi‑view inputs symmetrically. Trained on the large Laval Objaverse Dataset, the model achieves state‑of‑the‑art visual and photorealistic relighting quality, and demonstrates strong zero‑shot generalization across various relighting tasks.
By Hejun Wang, Jinxi Li, Junwei Jiang, Shiwei Mao, Hu Cheng, Shouwang Huang, Bo Yang
Recent advances in neural scene representations enable photorealistic novel-view synthesis, yet most methods remain tightly coupled to a single rendering paradigm, limiting their versatility and integration with conventional graphics workflows. We introduce Floating Radiance Networks (FlaRe), a neural scene representation combining explicit ray-traceable geometry with continuous neural radiance functions.
arXiv:2604. 05182v2 Announce Type: replace-cross Abstract: We introduce the Large Sparse Reconstruction Model to study how scaling transformer context windows affects feed-forward 3D reconstruction.
By Zhengqin Li, Cheng Zhang, Jakob Engel, Zhao Dong
DiffusionShadow introduces a diffusion-based shadow caching framework for neural volume rendering, compressing many pre‑computed shadow INRs into a single diffusion model conditioned on lighting direction. The method encodes shadow coefficient volumes as shadow INRs, trains the diffusion model to predict shadow INR weights at inference, and integrates directly with standard INR renderers without extra runtime sampling. Experiments demonstrate faster rendering than traditional approaches while avoiding the large storage overhead of independent INRs, producing shadows that closely match reference results.
By Kai-Chen Tung, Qi Wu, David Bauer, Mengjiao Han, Silvio Rizzi, Kwan-Liu Ma
arXiv:2502. 07531v5 Announce Type: replace-cross Abstract: Controllable image-to-video (I2V) generation transforms a reference image into a coherent video guided by user-specified control signals.
By Sixiao Zheng, Zimian Peng, Yanpeng Zhou, Yi Zhu, Hang Xu, Xiangru Huang, Yanwei Fu
The paper introduces a method that applies physically based rendering (PBR) within the latent space of variational autoencoders used in image diffusion models. By modifying the rendering equation and using a differentiable renderer, the authors can generate latent maps that guide content creation with physically accurate lighting. The approach is trained on a single rendered image and then shown to generalize to changes in scene geometry, lighting, and camera viewpoint.
By Vuk Radovanovic, Vishesh Gupta, Adrien Gruson, Binh-Son Hua
arXiv:2609.23733v1 Announce Type: new
Abstract: Feed-forward visual geometry models such as the Visual Geometry Grounded Transformer (VGGT) have recently enabled direct 3D reconstruction from multi-v...
By Abteen Arab, Guile Wu, Chengjie Huang, Dongfeng Bai
arXiv:2606. 22574v2 Announce Type: replace-cross Abstract: While synthetic data generation resolves the manual labeling bottleneck in computer vision, minimizing the syn-to-real domain gap requires optimizing rendering variables.
By Hooman Tavakoli Ghinani, Tatjana Legler, Martin Ruskowski
Reconstructing 3D shapes from a single image remains a fundamental yet challenging problem in computer vision. Traditional monocular 3D generation pipelines typically synthesize multiple views from a single input image before applying Neural Radiance Field (NeRF)-based reconstruction.
arXiv:2609.01306v1 Announce Type: cross
Abstract: Triangle-based neural rendering bridges neural scene representations and conventional graphics pipelines by optimizing explicit geometric primitives...
By Kaixuan Zhang, Minxian Li, Mingwu Ren, Xiatian Zhu