arXiv Computer Vision By Vitor Miguel Xavier Peres, Lara Volpato, Gabriel Ferri Scnheider, Soraia Raupp Musse

Emotion Intensity Matters: Generating Realistic Expressions in Virtual Humans with CVAEs

Read the original on arXiv Computer Vision →

This paper introduces a Conditional Variational Autoencoder (CVAEs) approach that generates realistic, controllable emotional facial expressions for virtual humans. Trained on a small dataset of 7,680 samples covering six basic emotions at low and high intensity, the model learns latent representations that preserve key expressive characteristics across intensity levels. The method enables animators to produce emotionally expressive virtual characters without actor performances or manual artistic effort.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Computation and Language
Sep 4

Chehre: An Emoji-Prompted Dataset to Explore Perceptual Flexibility in Video Language Models

Chehre is an emoji‑prompted video dataset designed to study perceptual flexibility in video language models. It contains 2,111 videos of 203 participants expressing 40 facial emojis, with each video annotated by about 30 perceivers, yielding 1,242 annotators in total. The dataset introduces a new task—distributional expression recognition—that evaluates a model’s ability to reproduce the variation seen in human annotations, and shows that persona prompting can shift model perception to better match human variability.

By Bita Azari, Zoe Stanley, Avneet Batra, Poorvi Bhatia, Hali Kil, Manolis Savva, Angelica Lim
arXiv Computer Vision
6d ago

Capturing Dynamics: The 4D Facial Expression Intensity Dataset

The paper introduces the 4D Facial Expression Intensity Dataset (4DFEID), comprising 2,869 mesh sequences that capture 3D, temporally continuous facial expressions with varied peak intensities and identities. Subjective intensity ratings were collected via crowdsourcing, yielding over 90,000 Likert-scale annotations. Baseline experiments show that spatial‑temporal graph models outperform traditional frame‑aggregation methods, highlighting the dataset’s value for dynamic 3D expression analysis.

By Zesheng Wang, Alexandre Bruckert, Pierre Lebreton, Patrick Le Callet, Yante Li, Guoying Zhao