Lightstage facial capture produces production-quality digital humans, but it is resource and labor-intensive. Multi-camera setups, hours of computation, and massive data storage create bottlenecks tha...
arXiv:2608.22655v1 Announce Type: new
Abstract: Creating re-topologized 3D facial meshes is essential for high-quality facial animation but remains labor-intensive and time-consuming. This dissertati...
By Xiang Li
Face Video Restoration (FVR) aims to recover high-fidelity facial videos from degraded input while preserving identity and semantic consistency across frames. Existing methods often struggle to simultaneously address three key challenges: identity shift, viewpoint-entangled guidance, and perceptual realism.
arXiv:2609.24158v1 Announce Type: new
Abstract: Reconstructing expressive and relightable 3D head avatars from monocular videos remains challenging in computer vision, as it requires accurate modelin...
By Jiankuo Zhao, Xiangyu Zhu, Jijie Li, Baiqin Wang, Shukai Chen, Zhen Lei
arXiv:2608.23410v1 Announce Type: new
Abstract: Photorealistic novel view synthesis of people remains challenging at high spatial resolutions and across multiple target cameras, where preserving iden...
By Federico Stella, Fei Jiang, Zhongshi Jiang, Zohar Barzelay, Emanuel Garbin, Amin Jourabloo, Liuhao Ge
arXiv:2606. 30347v1 Announce Type: cross Abstract: We present FFAvatar, a Transformer-based 3D Gaussian framework for fast construction of high-quality and animatable 4D head avatars from one or more reference portrait images.
By Jianjiang Yao, Ke Xian, Renxiang Dai, Robert Caiming Qiu
arXiv:2609.15032v1 Announce Type: new
Abstract: Live free-viewpoint visualization of real humans is critical for immersive communication and interactive digital experiences. Existing methods either r...
By Hanzhang Tu, Zhanfeng Liao, Wei Min, Jiajun Zhang, Yebin Liu
arXiv:2605. 04035v3 Announce Type: replace-cross Abstract: We propose HeadsUp, a scalable feed-forward method for reconstructing high-quality 3D Gaussian heads from large-scale multi-camera setups.
By Evangelos Ntavelis, Sean Wu, Mohamad Shahbazi, Fabio Maninchedda, Dmitry Kostiaev, Artem Sevastopolsky, Vittorio Megaro, Trevor Phillips, Alejandro Blumentals, Shridhar Ravikumar, Mehak Gupta, Reinhard Knothe, Jeronimo Bayer, Matthias Vestner, Simon Schaefer, Thomas Etterlin, Christian Zimmermann, Alexey Artemov, Mathias Deschler, Peter Kaufmann, Stefan Brugger, Sebastian Martin, Brian Amberg, Tom Runia
arXiv:2607. 16287v1 Announce Type: cross Abstract: Neural Radiance Fields (NeRF) have enabled photorealistic novel-view synthesis of 3D scenes and, in the facial domain, have been extended to reconstruct and animate 3D face models from a small number of images.
By Minh Tran
GenStream is a semantic streaming framework that replaces dense video frames with compact metadata—skeletal keypoints, camera parameters, and a static 3D background model—to enable generative reconstruction of human figures on the client side. By transmitting only structured information rather than full pixel data, it achieves over 99.9% bandwidth reduction compared to HEVC, as demonstrated on Olympic figure skating footage. The approach shifts computational load to the client and opens possibilities for volumetric avatar synthesis, multi‑view actor fusion, and personalized viewing experiences in a post‑codec era.
By Emanuele Artioli, Daniele Lorenzi, Shivi Vats, Farzad Tashtarian, Christian Timmerer
The paper introduces Temporal Residual Neural Radiance Fields for reconstructing dynamic human bodies from monocular video. It builds a temporal residual field independent of MLPs, reduces trainable parameters, speeds up rendering, and employs a multi‑dimensional loss to improve pixel‑level accuracy. Experiments show higher PSNR and SSIM than recent methods while being roughly 780 times faster than Anim‑NeRF and Neural Body.
By Tianle Du, Jie Wang, Xiaolong Xie, Wei Li, Pengxiang Su, Jie Liu
Projector-camera (ProCams) systems achieve active scene perception and controllable appearance manipulation via structured illumination, serving as a core infrastructure for spatial augmented reality, projection mapping, and surface reflectance acquisition. Existing inverse-rendering methods for ProCams deliver high-fidelity results but rely on time-consuming per-scene optimization, while mainstream feed-forward 3D reconstruction models produce baked appearance that cannot adapt to spatially varying projector illumination.