Generative Learning as a Tool to Improve Perception of Emotional Body Motion Expressions
arXiv:2606. 28769v1 Announce Type: new Abstract: Emotional body motion expressions are an essential element of non-verbal communication.
This paper introduces a Conditional Variational Autoencoder (CVAEs) approach that generates realistic, controllable emotional facial expressions for virtual humans. Trained on a small dataset of 7,680 samples covering six basic emotions at low and high intensity, the model learns latent representations that preserve key expressive characteristics across intensity levels. The method enables animators to produce emotionally expressive virtual characters without actor performances or manual artistic effort.
arXiv:2606. 28769v1 Announce Type: new Abstract: Emotional body motion expressions are an essential element of non-verbal communication.
Chehre is an emoji‑prompted video dataset designed to study perceptual flexibility in video language models. It contains 2,111 videos of 203 participants expressing 40 facial emojis, with each video annotated by about 30 perceivers, yielding 1,242 annotators in total. The dataset introduces a new task—distributional expression recognition—that evaluates a model’s ability to reproduce the variation seen in human annotations, and shows that persona prompting can shift model perception to better match human variability.
arXiv:2608. 15110v1 Announce Type: cross Abstract: Emotional 3D talking head generation aims to synthesize expressive facial animations with accurate lip synchronization.
Emotional 3D talking head generation aims to synthesize expressive facial animations with accurate lip synchronization. However, existing methods often rely on discrete emotion categories, which fail...
The paper introduces the 4D Facial Expression Intensity Dataset (4DFEID), comprising 2,869 mesh sequences that capture 3D, temporally continuous facial expressions with varied peak intensities and identities. Subjective intensity ratings were collected via crowdsourcing, yielding over 90,000 Likert-scale annotations. Baseline experiments show that spatial‑temporal graph models outperform traditional frame‑aggregation methods, highlighting the dataset’s value for dynamic 3D expression analysis.
arXiv:2609.13264v1 Announce Type: cross Abstract: Generating human-centric videos that preserve both visual identity and person-specific expressive behavior remains a fundamental challenge. In additi...
EMODY Flow is a lightweight flow‑matching framework that generates synchronized full‑body motion and facial expressions conditioned on speech and emotion. It attaches to a frozen Qwen‑3 Omni model, reusing its audio codecs to drive two parallel DiT generators for SMPL‑X body pose and FLAME facial expressions. An auxiliary emotion classifier at training time restores emotion sensitivity, enabling EMODY Flow to achieve state‑of‑the‑art gesture quality on BEAT2 and zero‑shot facial animation on TFHP, with significant improvements in FGD, Beat Correlation, and Diversity metrics.
arXiv:2609.24215v1 Announce Type: new Abstract: Although text-to-image models can accurately depict subjects and scenes, creators still struggle to specify the fine-grained emotions an image should c...
arXiv:2604. 15280v2 Announce Type: replace-cross Abstract: Understanding emotions is a fundamental ability for intelligent systems to be able to interact with humans.
Audio-driven talking head synthesis has achieved impressive progress in lip synchronization and visual quality, yet generating expressive emotional avatars with controllable intensity remains challenging, especially under real-time constraints. In this paper, we present GaussianEmoTalker, an audio-driven framework for real-time emotional talking head synthesis based on 3D Gaussian Splatting.
arXiv:2601. 00664v2 Announce Type: replace-cross Abstract: Talking head generation creates lifelike avatars from static portraits for virtual communication and content creation.
arXiv:2610.11023v1 Announce Type: new Abstract: Identity-preserving video generation aims to maintain a subject's identity while synthesizing realistic videos. Yet a single reference portrait capture...