AGORA is a new framework that extends 3D Gaussian Splatting with a generative adversarial network to produce high‑fidelity, animatable 3D head avatars. It introduces a lightweight FLAME‑conditioned deformation branch that predicts per‑Gaussian residuals for identity‑preserving, fine‑grained expression control, and a dual‑discriminator training scheme that enforces expression fidelity. The system achieves real‑time inference at 250 FPS on a single GPU and, for the first time, CPU‑only animatable 3DGS avatar synthesis at ~9 FPS.
By Ramazan Fazylov, Sergey Zagoruyko, Aleksandr Parkin, Stamatis Lefkimmiatis, Ivan Laptev
ChromaGS is a real‑time, language‑guided method for color editing of animatable 3D Gaussian head avatars. It augments each Gaussian primitive with learned soft assignments to semantic regions and decomposes colors into region‑level base colors and Gaussian‑level residuals, enabling coherent color transfer while preserving fine details. A two‑stage language pipeline translates natural‑language instructions into target colors, supporting both absolute and relative adjustments without requiring retraining.
By Antonio Canela, Jordi S\`anchez-Riera
arXiv:2609.38343v1 Announce Type: new
Abstract: We present SInGA, a novel method for learning Semantic Inpainting for animatable Gaussian head Avatars from a single image. Existing avatar approaches...
By Pilseo Park, Fizza Rubab, Yiying Tong
EmbedTalk introduces per‑Gaussian embeddings to drive speech‑driven facial deformations in real‑time talking head synthesis, replacing traditional tri‑plane encodings. This approach improves rendering quality, lip synchronisation, and motion consistency compared to prior 3D Gaussian Splatting methods while producing more compact models that run at 60+ FPS on a laptop GPU. The technique demonstrates competitive performance against state‑of‑the‑art generative models.
By Arpita Saggar, Jonathan C. Darling, Duygu Sarikaya, David C. Hogg
arXiv:2608. 19900v1 Announce Type: new Abstract: For full-body avatars, modeling surface dynamics is crucial for overcoming the uncanny valley and achieving perceptual realism.
By Guoxing Sun, Heming Zhu, Linjie Lyu, Pascal Fua, Christian Theobalt, Marc Habermann
For full-body avatars, modeling surface dynamics is crucial for overcoming the uncanny valley and achieving perceptual realism. Person-agnostic methods recover static 3D avatars from monocular images, videos, or text prompts, but their skeleton-driven animations lack realistic surface dynamics such as clothing wrinkles.
arXiv:2608.21136v1 Announce Type: new
Abstract: Recently, open-vocabulary zero-shot 3D scene understanding using vision foundation models has emerged as a promising alternative to data-intensive supe...
By Jie Xu, Na Zhao
arXiv:2609.12850v1 Announce Type: new
Abstract: Accurate head modeling requires a stable yet expressive geometric representation. Existing Gaussian-based head avatars commonly rely on parametric temp...
By Lei Shi, Sen Peng, Zhiyang Deng, Zhonggui Chen, Xiaohu Guo, Baorong Yang, Xiao Dong
arXiv:2606. 30347v1 Announce Type: cross Abstract: We present FFAvatar, a Transformer-based 3D Gaussian framework for fast construction of high-quality and animatable 4D head avatars from one or more reference portrait images.
By Jianjiang Yao, Ke Xian, Renxiang Dai, Robert Caiming Qiu
3D Gaussian Splatting has achieved remarkable success in photorealistic and efficient rendering, leading to a rapid increase in 3D assets represented by 3D Gaussian primitives. Directly rigging these assets with arbitrary skeleton topologies is highly desirable.
arXiv:2507. 11061v3 Announce Type: replace-cross Abstract: Recent advances in 3D neural representations and instance-level editing models have enabled the efficient creation of high-quality 3D content.
By Hayeon Kim, Ji Ha Jang, Se Young Chun
WilLaGS introduces a unified framework that enhances 3D Gaussian Splatting for in-the-wild scenes by learning a continuous global appearance manifold with a β‑VAE and generating dynamic Tri‑Plane features for spatially‑varying local illumination. It also incorporates a self‑supervised perceptual masking mechanism using a Teacher‑Student EMA architecture to suppress transient artifacts and identify inconsistent regions. Experiments on multiple datasets show that WilLaGS achieves state‑of‑the‑art reconstruction quality and novel view synthesis while preserving real‑time rendering efficiency.
By Yuhao Bai, Qianqiu Tan, Lilong Chen, Huanhuan Lv, Lijun Chen