arXiv Machine Learning By Huakun Liu, Miao Cheng, Xin Wei, Felix Dollack, Victor Schneider, Hideaki Uchiyama, Chia-huei Tseng, Yoshifumi Kitamura, Monica Perusquia-Hernandez

Generative Learning as a Tool to Improve Perception of Emotional Body Motion Expressions

Read the original on arXiv Machine Learning →

arXiv:2606. 28769v1 Announce Type: new Abstract: Emotional body motion expressions are an essential element of non-verbal communication.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 16

EMODY Flow: Emotion-Aware Audio-Driven Full-Body Motion Generation

EMODY Flow is a lightweight flow‑matching framework that generates synchronized full‑body motion and facial expressions conditioned on speech and emotion. It attaches to a frozen Qwen‑3 Omni model, reusing its audio codecs to drive two parallel DiT generators for SMPL‑X body pose and FLAME facial expressions. An auxiliary emotion classifier at training time restores emotion sensitivity, enabling EMODY Flow to achieve state‑of‑the‑art gesture quality on BEAT2 and zero‑shot facial animation on TFHP, with significant improvements in FGD, Beat Correlation, and Diversity metrics.

By Harsh Kumar Agarwal, Xavier Alameda-Pineda, Olivier Perrotin
arXiv Computer Vision
Aug 25

Emotion Intensity Matters: Generating Realistic Expressions in Virtual Humans with CVAEs

This paper introduces a Conditional Variational Autoencoder (CVAEs) approach that generates realistic, controllable emotional facial expressions for virtual humans. Trained on a small dataset of 7,680 samples covering six basic emotions at low and high intensity, the model learns latent representations that preserve key expressive characteristics across intensity levels. The method enables animators to produce emotionally expressive virtual characters without actor performances or manual artistic effort.

By Vitor Miguel Xavier Peres, Lara Volpato, Gabriel Ferri Scnheider, Soraia Raupp Musse
arXiv Computer Vision
Aug 28

HUG-VIS: A Multimodal Benchmark for Human-centered Understanding and Generation in Visual Intelligence

HUG‑VIS is a unified multimodal benchmark for human‑centered visual intelligence, comprising 8,400 half‑body videos of 30 professional actors performing 280 emotion‑action prompts in Mandarin. The dataset provides synchronized video, audio, text, and alpha mattes for four tasks—human emotion recognition, video generation, voice cloning, and video matting—allowing evaluation of both open‑ and closed‑source models under a zero‑shot protocol. Results reveal that linguistic cues dominate emotion recognition, visual affect is weakest, and that automatic metrics and human judgments diverge in generation and cloning tasks, while motion‑related boundary fidelity remains a key challenge for matting.

By Fei Ma, Zebang Cheng, Minghui Li, Hongbo Xu, Yuyong Tan, Yihua Shao, Hanling Wang, Zhou Liu, Yuqing Gao, Dong Wang, Long Ma, Laizhong Cui, Nicu Sebe, Qi Tian