arXiv:2609.24158v1 Announce Type: new
Abstract: Reconstructing expressive and relightable 3D head avatars from monocular videos remains challenging in computer vision, as it requires accurate modelin...
By Jiankuo Zhao, Xiangyu Zhu, Jijie Li, Baiqin Wang, Shukai Chen, Zhen Lei
arXiv:2608. 19900v1 Announce Type: new Abstract: For full-body avatars, modeling surface dynamics is crucial for overcoming the uncanny valley and achieving perceptual realism.
By Guoxing Sun, Heming Zhu, Linjie Lyu, Pascal Fua, Christian Theobalt, Marc Habermann
For full-body avatars, modeling surface dynamics is crucial for overcoming the uncanny valley and achieving perceptual realism. Person-agnostic methods recover static 3D avatars from monocular images, videos, or text prompts, but their skeleton-driven animations lack realistic surface dynamics such as clothing wrinkles.
arXiv:2608.23549v1 Announce Type: new
Abstract: Rendering views using 3D scene representations such as Gaussian Splatting (3DGS), Neural Radiance Fields (NeRF), meshes, or even point clouds produces...
By Khiem Vuong, Deva Ramanan, Srinivasa Narasimhan
arXiv:2512.03593v2 Announce Type: replace
Abstract: We present a CloseUpAvatar - a novel approach for articulated human avatar representation supporting a wider range of camera motions, while preserv...
By David Svitov, Pietro Morerio, Lourdes Agapito, Alessio Del Bue
Rendering views using 3D scene representations such as Gaussian Splatting (3DGS), Neural Radiance Fields (NeRF), meshes, or even point clouds produces artifacts when input views are sparse or target v...
arXiv:2610.02207v1 Announce Type: cross
Abstract: 3D Gaussian avatars support fast rendering, however, their real-time animation is often challenged by the costly neural inference. We address this bo...
By Ramazan Fazylov, Stamatis Lefkimmiatis, Ivan Laptev
arXiv:2608.29925v1 Announce Type: new
Abstract: Controllable image relighting is an important problem in image editing, and hand-drawn scribbles provide an intuitive interface for specifying the desi...
By Xuanpu Zhang, Xuesong Niu, Haoxiang Cao, Ruidong Chen, Jianhao Zeng, Changqian Yu
Projector-camera (ProCams) systems achieve active scene perception and controllable appearance manipulation via structured illumination, serving as a core infrastructure for spatial augmented reality, projection mapping, and surface reflectance acquisition. Existing inverse-rendering methods for ProCams deliver high-fidelity results but rely on time-consuming per-scene optimization, while mainstream feed-forward 3D reconstruction models produce baked appearance that cannot adapt to spatially varying projector illumination.
RelightFormer is a feed‑forward generative Transformer that performs single‑ and multi‑view image relighting without explicit intrinsic property estimation. It incorporates a latent illumination module that injects target environment maps into spatial features via cross‑attention, and uses permutation‑invariant positional encodings to process unordered multi‑view inputs symmetrically. Trained on the large Laval Objaverse Dataset, the model achieves state‑of‑the‑art visual and photorealistic relighting quality, and demonstrates strong zero‑shot generalization across various relighting tasks.
By Hejun Wang, Jinxi Li, Junwei Jiang, Shiwei Mao, Hu Cheng, Shouwang Huang, Bo Yang
arXiv:2607. 22830v2 Announce Type: replace Abstract: In visual storytelling, human performances are central to creative intent and narrative meaning.
By Yuancheng Xu, Mingming He, Pablo Salamanca, Li Ma, Yash Kant, Emmett Steven, Paul Debevec, Ning Yu
arXiv:2606.24232v2 Announce Type: replace
Abstract: We introduce FiCA, a Feed-forward, instant Gaussian Codec Avatar generation pipeline that creates lifelike avatars from a single portrait image. Ge...
By Kim Youwang, Zhengyu Yang, Liuhao Ge, Yu Rong, Timur Bagautdinov, Su Zhaoen, Nir Sopher, Jovan Popovi\'c, Teng Deng, Tae-Hyun Oh, Chen Cao