The paper introduces GALA, a distillation technique that replaces costly neural decoding in 3D Gaussian avatars with a shallow MLP predicting blendshape coefficients, enabling real‑time animation. By constructing a basis via block‑local PCA under a rendering‑aware metric, GALA achieves high fidelity while reducing memory usage. Experiments on three avatar models show up to three orders of magnitude lower CPU cost and frame rates up to 60fps on mobile devices.
By Ramazan Fazylov, Stamatis Lefkimmiatis, Ivan Laptev
AGORA is a new framework that extends 3D Gaussian Splatting with a generative adversarial network to produce high‑fidelity, animatable 3D head avatars. It introduces a lightweight FLAME‑conditioned deformation branch that predicts per‑Gaussian residuals for identity‑preserving, fine‑grained expression control, and a dual‑discriminator training scheme that enforces expression fidelity. The system achieves real‑time inference at 250 FPS on a single GPU and, for the first time, CPU‑only animatable 3DGS avatar synthesis at ~9 FPS.
By Ramazan Fazylov, Sergey Zagoruyko, Aleksandr Parkin, Stamatis Lefkimmiatis, Ivan Laptev
arXiv:2609.12850v1 Announce Type: new
Abstract: Accurate head modeling requires a stable yet expressive geometric representation. Existing Gaussian-based head avatars commonly rely on parametric temp...
By Lei Shi, Sen Peng, Zhiyang Deng, Zhonggui Chen, Xiaohu Guo, Baorong Yang, Xiao Dong
arXiv:2608. 19900v1 Announce Type: new Abstract: For full-body avatars, modeling surface dynamics is crucial for overcoming the uncanny valley and achieving perceptual realism.
By Guoxing Sun, Heming Zhu, Linjie Lyu, Pascal Fua, Christian Theobalt, Marc Habermann
The paper introduces DirectSwap, a mask‑free video head‑swapping method that leverages a newly created cross‑identity paired dataset, HeadSwapBench. By synthesizing expression‑synchronized video pairs from real footage, the authors provide frame‑aligned ground truth for full‑reference evaluation of identity, expression, pose, reconstruction fidelity, and temporal stability. DirectSwap outperforms traditional same‑identity masked reconstruction, especially when head silhouettes change, and can restore non‑head content without external segmentation.
By Yanan Wang, Shengcai Liao, Panwen Hu, Xin Li, Fan Yang, Guangxi Liu, Xiaodan Liang
For full-body avatars, modeling surface dynamics is crucial for overcoming the uncanny valley and achieving perceptual realism. Person-agnostic methods recover static 3D avatars from monocular images, videos, or text prompts, but their skeleton-driven animations lack realistic surface dynamics such as clothing wrinkles.
arXiv:2606. 30347v1 Announce Type: cross Abstract: We present FFAvatar, a Transformer-based 3D Gaussian framework for fast construction of high-quality and animatable 4D head avatars from one or more reference portrait images.
By Jianjiang Yao, Ke Xian, Renxiang Dai, Robert Caiming Qiu
Facial movements convey subtle and important information that is critical for human social communication. Optical methods for face capture are difficult or impossible to use when the face is occluded...
arXiv:2609.38343v1 Announce Type: new
Abstract: We present SInGA, a novel method for learning Semantic Inpainting for animatable Gaussian head Avatars from a single image. Existing avatar approaches...
By Pilseo Park, Fizza Rubab, Yiying Tong
arXiv:2609.24158v1 Announce Type: new
Abstract: Reconstructing expressive and relightable 3D head avatars from monocular videos remains challenging in computer vision, as it requires accurate modelin...
By Jiankuo Zhao, Xiangyu Zhu, Jijie Li, Baiqin Wang, Shukai Chen, Zhen Lei
arXiv:2607. 22830v2 Announce Type: replace Abstract: In visual storytelling, human performances are central to creative intent and narrative meaning.
By Yuancheng Xu, Mingming He, Pablo Salamanca, Li Ma, Yash Kant, Emmett Steven, Paul Debevec, Ning Yu
The paper presents emg2face, a system that uses high‑density surface electromyography (HD‑sEMG) to capture facial expressions without optical cameras, addressing issues of occlusion and privacy. It records 64 EMG channels with textile grids, synchronizes the data with video using analog audio bursts, and fits a high‑resolution parametric head model to 3D facial landmarks. A deep neural network then predicts blendshape parameters from the EMG signals, enabling real‑time facial animation on various characters.
By Ganidhu Abey, Wendy Greening, Ashika Kamboj, Leonhard Helminger, Abhijeet Ghosh, Karel Petranek, Sergio Orts Escolano, Dinesh K. Pai