arXiv Computer Vision

FaceSnap: Real-Time Personalized Lightstage Facial Performance Capture

arXiv Machine Learning
Jul 2

Large-Scale High-Quality 3D Gaussian Head Reconstruction from Multi-View Captures

arXiv:2605. 04035v3 Announce Type: replace-cross Abstract: We propose HeadsUp, a scalable feed-forward method for reconstructing high-quality 3D Gaussian heads from large-scale multi-camera setups.

By Evangelos Ntavelis, Sean Wu, Mohamad Shahbazi, Fabio Maninchedda, Dmitry Kostiaev, Artem Sevastopolsky, Vittorio Megaro, Trevor Phillips, Alejandro Blumentals, Shridhar Ravikumar, Mehak Gupta, Reinhard Knothe, Jeronimo Bayer, Matthias Vestner, Simon Schaefer, Thomas Etterlin, Christian Zimmermann, Alexey Artemov, Mathias Deschler, Peter Kaufmann, Stefan Brugger, Sebastian Martin, Brian Amberg, Tom Runia
arXiv AI
Sep 17

GenStream: Semantic Streaming Framework for Generative Reconstruction of Human-centric Media

GenStream is a semantic streaming framework that replaces dense video frames with compact metadata—skeletal keypoints, camera parameters, and a static 3D background model—to enable generative reconstruction of human figures on the client side. By transmitting only structured information rather than full pixel data, it achieves over 99.9% bandwidth reduction compared to HEVC, as demonstrated on Olympic figure skating footage. The approach shifts computational load to the client and opens possibilities for volumetric avatar synthesis, multi‑view actor fusion, and personalized viewing experiences in a post‑codec era.

By Emanuele Artioli, Daniele Lorenzi, Shivi Vats, Farzad Tashtarian, Christian Timmerer
arXiv Computer Vision
Sep 7

Temporal Residual Neural Radiance Fields for Monocular Video Dynamic Human Body Reconstruction

The paper introduces Temporal Residual Neural Radiance Fields for reconstructing dynamic human bodies from monocular video. It builds a temporal residual field independent of MLPs, reduces trainable parameters, speeds up rendering, and employs a multi‑dimensional loss to improve pixel‑level accuracy. Experiments show higher PSNR and SSIM than recent methods while being roughly 780 times faster than Anim‑NeRF and Neural Body.

By Tianle Du, Jie Wang, Xiaolong Xie, Wei Li, Pengxiang Su, Jie Liu
Hugging Face Trending Papers
Jul 20

FF-ProCams: Feed-Forward Gaussian Splatting for Projector-Camera System

Projector-camera (ProCams) systems achieve active scene perception and controllable appearance manipulation via structured illumination, serving as a core infrastructure for spatial augmented reality, projection mapping, and surface reflectance acquisition. Existing inverse-rendering methods for ProCams deliver high-fidelity results but rely on time-consuming per-scene optimization, while mainstream feed-forward 3D reconstruction models produce baked appearance that cannot adapt to spatially varying projector illumination.