Hugging Face Trending Papers

FacePlex: Full-Duplex Joint Speech-Facial Motion Generation for Conversational Avatars

Read the original on Hugging Face Trending Papers →

Natural face-to-face conversation requires real-time speech generation together with synchronized facial motion. Existing systems only partially address this problem: speech-only full-duplex models can generate speech in real time but do not produce facial motion, while audio-driven facial motion models animate a face from already available audio rather than jointly generating speech and motion online.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv Machine Learning
Sep 22

AVTR-1: Open Stack for Real-Time Interactive Avatars

arXiv:2609.22913v1 Announce Type: cross Abstract: Talking-head and dyadic models now achieve real-time inference, yet fast motion generation alone does not produce an interactive conversation. A live...

By Artem Kravtsov, Dmitrii Ziganshin, Vsevolod Poletaev, Gleb Balitskiy, Anastasia Tikhonova, Egor Burkov, Vadim Lebedev
arXiv Computer Vision
Sep 21

Personalizing Causal Audio-Driven Facial Motion via Dynamic Multi-modal Retrieval

The paper introduces an end‑to‑end framework for personalized audio‑driven facial motion that operates in real time without look‑ahead. It combines a causal multi‑resolution motion tokenizer, which captures both global temporal context and fine articulatory details, with a multi‑modal style retriever that pulls stylistic priors from arbitrary reference footage using ongoing audio and motion queries. This approach allows high‑fidelity, identity‑consistent animation from just a few casually recorded clips, outperforming existing methods in lip‑sync accuracy, identity consistency, and perceived realism while maintaining real‑time streaming constraints.

By Xuangeng Chu, Yu Han, Wei Mao, Shih-En Wei
arXiv Computer Vision
Sep 21

GestureFAR: Streaming Co-Speech Gesture Generation with Flow Autoregression

GestureFAR is a flow‑autoregressive framework that generates natural co‑speech gestures from streaming speech while preserving causality and continuous motion expressiveness. It autoregresses over continuous motion latents using a transformer for audio‑motion context and a flow‑matching head to sample the next latent. A head‑only flow distillation strategy further reduces latency by collapsing multi‑step flow sampling into a single network evaluation, enabling real‑time token‑causal generation with improved quality‑latency trade‑off on the BEAT2 benchmark.

By Pinxin Liu, Haiyang Liu, Jiahao Luo, Junhua Huang, Chunhao Zou, Luchuan Song