arXiv Machine Learning

GHARP: Real-time Gaussian Head Animation from Large-scale Reconstruction Prior

arXiv AI
Oct 2

One Basis to Animate Them All: Gaussian Blendshape Distillation for Real-Time Avatars

The paper introduces GALA, a distillation technique that replaces costly neural decoding in 3D Gaussian avatars with a shallow MLP predicting blendshape coefficients, enabling real‑time animation. By constructing a basis via block‑local PCA under a rendering‑aware metric, GALA achieves high fidelity while reducing memory usage. Experiments on three avatar models show up to three orders of magnitude lower CPU cost and frame rates up to 60fps on mobile devices.

By Ramazan Fazylov, Stamatis Lefkimmiatis, Ivan Laptev
arXiv Computer Vision
Sep 18

AGORA: Adversarial Generation Of Real-time Animatable 3D Gaussian Head Avatars

AGORA is a new framework that extends 3D Gaussian Splatting with a generative adversarial network to produce high‑fidelity, animatable 3D head avatars. It introduces a lightweight FLAME‑conditioned deformation branch that predicts per‑Gaussian residuals for identity‑preserving, fine‑grained expression control, and a dual‑discriminator training scheme that enforces expression fidelity. The system achieves real‑time inference at 250 FPS on a single GPU and, for the first time, CPU‑only animatable 3DGS avatar synthesis at ~9 FPS.

By Ramazan Fazylov, Sergey Zagoruyko, Aleksandr Parkin, Stamatis Lefkimmiatis, Ivan Laptev
arXiv Computer Vision
Sep 11

DirectSwap: Paired, Mask-Free Video Head Swapping with Full-Reference Evaluation

The paper introduces DirectSwap, a mask‑free video head‑swapping method that leverages a newly created cross‑identity paired dataset, HeadSwapBench. By synthesizing expression‑synchronized video pairs from real footage, the authors provide frame‑aligned ground truth for full‑reference evaluation of identity, expression, pose, reconstruction fidelity, and temporal stability. DirectSwap outperforms traditional same‑identity masked reconstruction, especially when head silhouettes change, and can restore non‑head content without external segmentation.

By Yanan Wang, Shengcai Liao, Panwen Hu, Xin Li, Fan Yang, Guangxi Liu, Xiaodan Liang
arXiv Computer Vision
Aug 21

ID-V2V: Identity-Preserving Video Restylization

arXiv:2607. 22830v2 Announce Type: replace Abstract: In visual storytelling, human performances are central to creative intent and narrative meaning.

By Yuancheng Xu, Mingming He, Pablo Salamanca, Li Ma, Yash Kant, Emmett Steven, Paul Debevec, Ning Yu
arXiv Machine Learning
3d ago

emg2face: Expressive Facial Animation with High-Density Surface EMG

The paper presents emg2face, a system that uses high‑density surface electromyography (HD‑sEMG) to capture facial expressions without optical cameras, addressing issues of occlusion and privacy. It records 64 EMG channels with textile grids, synchronizes the data with video using analog audio bursts, and fits a high‑resolution parametric head model to 3D facial landmarks. A deep neural network then predicts blendshape parameters from the EMG signals, enabling real‑time facial animation on various characters.

By Ganidhu Abey, Wendy Greening, Ashika Kamboj, Leonhard Helminger, Abhijeet Ghosh, Karel Petranek, Sergio Orts Escolano, Dinesh K. Pai