arXiv:2608. 15110v1 Announce Type: cross Abstract: Emotional 3D talking head generation aims to synthesize expressive facial animations with accurate lip synchronization.
By Peng Jia, Li Dai, Zhen Xiao, Xueliang Liu, Jia Li
Emotional 3D talking head generation aims to synthesize expressive facial animations with accurate lip synchronization. However, existing methods often rely on discrete emotion categories, which fail...
Audio-driven talking head synthesis has achieved impressive progress in lip synchronization and visual quality, yet generating expressive emotional avatars with controllable intensity remains challenging, especially under real-time constraints. In this paper, we present GaussianEmoTalker, an audio-driven framework for real-time emotional talking head synthesis based on 3D Gaussian Splatting.
arXiv:2503. 14295v3 Announce Type: replace-cross Abstract: Recent advancements in audio-driven talking face generation have made great progress in lip synchronization.
By Baiqin Wang, Xiangyu Zhu, Fan Shen, Hao Xu, Zhen Lei
arXiv:2606. 28568v1 Announce Type: cross Abstract: Speech-driven 3D facial animation methods face significant challenges in simultaneously achieving high-fidelity motion and precise artistic control at production quality.
By Arthur Josi, Emeline Got, Abdallah Dib, Luiz Gustavo Hafemann, Rafael M. O. Cruz
arXiv:2609.17422v1 Announce Type: new
Abstract: Audio-driven digital human generation plays an important role in virtual communication, immersive interaction, and media production. With the developme...
By Ziheng Yang, Yinfeng Yu, Yongming Li
The paper introduces Emo-DVS, a large-scale, multimodal dataset combining event camera, audio, and text data for emotion recognition, designed to mitigate privacy concerns associated with RGB cameras. It proposes the Information‑Guided Gated Fusion (IGF) framework, which pre‑trains an event encoder on the dataset’s FAU subset, adaptively gates modalities to reduce noise, and aligns cross‑modal representations via mutual information maximization. Experiments show that IGF outperforms existing methods on this challenging tri‑modal benchmark.
By Jiaqi Chen, Qinfu Xu, Hao Zhuang, Liyuan Pan
arXiv:2602. 07106v2 Announce Type: replace-cross Abstract: Omni-modal large language models (OLLMs) aim to unify multimodal understanding and generation, yet extending them to jointly produce speech and 3D facial animation remains largely unexplored despite its importance for natural human-computer interaction.
By Haoyu Zhang, Zhipeng Li, Yiwen Guo, Tianshu Yu
arXiv:2607. 16287v1 Announce Type: cross Abstract: Neural Radiance Fields (NeRF) have enabled photorealistic novel-view synthesis of 3D scenes and, in the facial domain, have been extended to reconstruct and animate 3D face models from a small number of images.
By Minh Tran
arXiv:2608. 05218v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) enables fast, photorealistic talking-head rendering, yet accurate lip articulation remains elusive: mouth motion is often over-smoothed and may violate hard articulatory constraints such as bilabial closures, producing the notorious ``leaky mouth'' artifact.
By Ao Fu, Yi Zhou
arXiv:2609.09924v1 Announce Type: new
Abstract: Emotion Recognition in Conversations (ERC) requires integrating heterogeneous textual, audio, and visual cues while accounting for conversational conte...
By Oriol Mar\'in, Roger Mar\'i, Gloria Haro, Rafael Redondo
Face Video Restoration (FVR) aims to recover high-fidelity facial videos from degraded input while preserving identity and semantic consistency across frames. Existing methods often struggle to simultaneously address three key challenges: identity shift, viewpoint-entangled guidance, and perceptual realism.