arXiv Machine Learning

emg2face: Expressive Facial Animation with High-Density Surface EMG

The paper presents emg2face, a system that uses high‑density surface electromyography (HD‑sEMG) to capture facial expressions without optical cameras, addressing issues of occlusion and privacy. It records 64 EMG channels with textile grids, synchronizes the data with video using analog audio bursts, and fits a high‑resolution parametric head model to 3D facial landmarks. A deep neural network then predicts blendshape parameters from the EMG signals, enabling real‑time facial animation on various characters.

arXiv Machine Learning
2d ago

GHARP: Real-time Gaussian Head Animation from Large-scale Reconstruction Prior

arXiv:2610.10945v1 Announce Type: cross Abstract: We present GHARP (Real-time Gaussian Head Animation from Large-scale Reconstruction Prior), a method that animates 3D human heads in real time from a...

By Ali Benlalah, Sepehr Johari, Patricia Vitoria, Armin Kappeler, Artem Sevastopolsky, Alexander Jung, Gabriele Fanelli, Kevin Mader, Manuel Breitenstein, Claudia Pl\"uss, Jan R\"uegg, Simon Biland, Thomas Etterlin, Dmitry Kostiaev, Mathias Deschler, Brian Amberg, Sebastian Martin
arXiv AI
Oct 2

One Basis to Animate Them All: Gaussian Blendshape Distillation for Real-Time Avatars

The paper introduces GALA, a distillation technique that replaces costly neural decoding in 3D Gaussian avatars with a shallow MLP predicting blendshape coefficients, enabling real‑time animation. By constructing a basis via block‑local PCA under a rendering‑aware metric, GALA achieves high fidelity while reducing memory usage. Experiments on three avatar models show up to three orders of magnitude lower CPU cost and frame rates up to 60fps on mobile devices.

By Ramazan Fazylov, Stamatis Lefkimmiatis, Ivan Laptev
arXiv Machine Learning
Sep 16

EMODY Flow: Emotion-Aware Audio-Driven Full-Body Motion Generation

EMODY Flow is a lightweight flow‑matching framework that generates synchronized full‑body motion and facial expressions conditioned on speech and emotion. It attaches to a frozen Qwen‑3 Omni model, reusing its audio codecs to drive two parallel DiT generators for SMPL‑X body pose and FLAME facial expressions. An auxiliary emotion classifier at training time restores emotion sensitivity, enabling EMODY Flow to achieve state‑of‑the‑art gesture quality on BEAT2 and zero‑shot facial animation on TFHP, with significant improvements in FGD, Beat Correlation, and Diversity metrics.

By Harsh Kumar Agarwal, Xavier Alameda-Pineda, Olivier Perrotin
arXiv Computer Vision
Sep 4

Occlusion-Robust Multimodal Emotion Recognition in VR via Fusion of Facial Images and EMG

The paper presents a method for emotion recognition in virtual reality where head‑mounted displays occlude the upper face. By fusing lower‑face video with electromyography (EMG) signals from the occluded upper face, the authors achieve a 51% macro‑F1 score across seven emotional categories, outperforming image‑only and EMG‑only baselines. A new synchronized multimodal dataset from 20 participants is introduced and will be shared under an ethical‑use agreement.

By Birgit Nierula, Karam Tomotaki-Dawoud, Mert Akguel, Mustafa Tevfik Lafci, David Przewozny, Anna Hilsmann, Peter Eisert, Sebastian Bosse
arXiv Computer Vision
Oct 1

MegaAvatar: Controllable Talking Avatar Generation

arXiv:2609.39273v1 Announce Type: new Abstract: This report presents \textbf{MegaAvatar}, a controllable talking avatar generation framework built on top of the Wan2.2-TI2V-5B model. Compared with pr...

By Junyao Gao, Sibo Liu, Weidong Zhang, Cairong Zhao, Jun Zhang
arXiv Computer Vision
6d ago

Capturing Dynamics: The 4D Facial Expression Intensity Dataset

The paper introduces the 4D Facial Expression Intensity Dataset (4DFEID), comprising 2,869 mesh sequences that capture 3D, temporally continuous facial expressions with varied peak intensities and identities. Subjective intensity ratings were collected via crowdsourcing, yielding over 90,000 Likert-scale annotations. Baseline experiments show that spatial‑temporal graph models outperform traditional frame‑aggregation methods, highlighting the dataset’s value for dynamic 3D expression analysis.

By Zesheng Wang, Alexandre Bruckert, Pierre Lebreton, Patrick Le Callet, Yante Li, Guoying Zhao