FacePlex: Full-Duplex Joint Speech-Facial Motion Generation for Conversational Avatars
arXiv:2606. 30145v1 Announce Type: new Abstract: Natural face-to-face conversation requires real-time speech generation together with synchronized facial motion.
arXiv:2606. 30145v1 Announce Type: new Abstract: Natural face-to-face conversation requires real-time speech generation together with synchronized facial motion.
Natural face-to-face conversation requires real-time speech generation together with synchronized facial motion. Existing systems only partially address this problem: speech-only full-duplex models can generate speech in real time but do not produce facial motion, while audio-driven facial motion models animate a face from already available audio rather than jointly generating speech and motion online.
Developers can now build fast speech-to-speech experiences into their applications
arXiv:2607. 03118v1 Announce Type: cross Abstract: We introduce Vidu S1, a real-time interactive video generation model supporting voice control of digital characters.
arXiv:2503. 14295v3 Announce Type: replace-cross Abstract: Recent advancements in audio-driven talking face generation have made great progress in lip synchronization.
arXiv:2609.22913v1 Announce Type: cross Abstract: Talking-head and dyadic models now achieve real-time inference, yet fast motion generation alone does not produce an interactive conversation. A live...
We’re releasing a more advanced speech-to-speech model and new API capabilities including MCP server support, image input, and SIP phone calling support.
arXiv:2609.39273v1 Announce Type: new Abstract: This report presents \textbf{MegaAvatar}, a controllable talking avatar generation framework built on top of the Wan2.2-TI2V-5B model. Compared with pr...
GestureFAR is a flow‑autoregressive framework that generates natural co‑speech gestures from streaming speech while preserving causality and continuous motion expressiveness. It autoregresses over continuous motion latents using a transformer for audio‑motion context and a flow‑matching head to sample the next latent. A head‑only flow distillation strategy further reduces latency by collapsing multi‑step flow sampling into a single network evaluation, enabling real‑time token‑causal generation with improved quality‑latency trade‑off on the BEAT2 benchmark.
GPT-Live enables continuous voice interaction with AI, using a turnless speech model and low-latency architecture for faster, more natural conversations.
Our latest voice model has improved precision and lower latency to make voice interactions more fluid, natural and precise.