Vidu S2 is a system that includes Vidu S2-Avatar, a real‑time interactive digital‑character model, and Vidu S2-Editing, a real‑time video editing model. It enables real‑time 720p video generation with dynamic references and improved instruction following, such as dancing, and allows real‑time editing of video streams for style rendering, clothing replacement, character replacement, and background replacement. Experiments show Vidu S2 outperforms all baselines, and a playable online demo is available at https://vidu.com/vidu-stream.
By Jintao Zhang, Kai Jiang, Jintao Chen, Xu Wang, Deyuan Liu, Jungang Li, Dechuang Chen, Ming Lin, Jingjiang Zhou, Haopeng Jin, Qi Jia, Xiaohang Wang, Yaole Wang, Zhanqiang Zhang, Ran Li, Zhengkun Huang, Shuyue Xiong, Yuji Wang, Zikun Dai, Hui He, Yang Luo, Mang Ning, Weiqi Feng, Chengyang Ye, Xinyue Lin, Min Zhao, Hongzhou Zhu, Hengkai Tan, Zeyuan Wang, Chendong Xiang, Kaiwen Zheng, Zhijie Deng, Fan Bao, Jianfei Chen, Jun Zhu
arXiv:2607.26694v3 Announce Type: replace
Abstract: We present Visko Orbis 1.0, a Live Model for real-time, interactive long video generation. Users can change the prompt at any moment during generat...
By Xiangbo Gao, Siyuan Yang, Ping He, Mingyang Wu, Yuheng Wu, Yushen Zuo, Jiongze Yu, Ryan Cui, Hongyuan Hua, Devin Ma, Xiao Jin, Yubo Ruan, Qing Yin, Jie Yang, Zhengzhong Tu
arXiv:2609.22913v1 Announce Type: cross
Abstract: Talking-head and dyadic models now achieve real-time inference, yet fast motion generation alone does not produce an interactive conversation. A live...
By Artem Kravtsov, Dmitrii Ziganshin, Vsevolod Poletaev, Gleb Balitskiy, Anastasia Tikhonova, Egor Burkov, Vadim Lebedev
arXiv:2609.39273v1 Announce Type: new
Abstract: This report presents \textbf{MegaAvatar}, a controllable talking avatar generation framework built on top of the Wan2.2-TI2V-5B model. Compared with pr...
By Junyao Gao, Sibo Liu, Weidong Zhang, Cairong Zhao, Jun Zhang
arXiv:2510. 00481v2 Announce Type: replace-cross Abstract: In 2025, Large Language Model (LLM) services have launched a new feature -- AI video chat -- allowing users to interact with AI agents via real-time video communication (RTC), just like chatting with real people.
By Jiayang Xu, Xiangjie Huang, Zijie Li, Antariksh Verma, Zili Meng
arXiv:2608. 11590v1 Announce Type: cross Abstract: Human voice generation has made rapid progress in speech generation, singing voice generation, voice cloning, and voice editing.
By Haowei Lou, Hye-Young Paik, Dai Jia, Kai Li, Lina Yao
arXiv:2601. 00664v2 Announce Type: replace-cross Abstract: Talking head generation creates lifelike avatars from static portraits for virtual communication and content creation.
By Taekyung Ki, Sangwon Jang, Jaehyeong Jo, Jaehong Yoon, Sung Ju Hwang
WanPE is a 397‑B parameter prompt‑enhancement model that learns director‑level cinematic planning from 1.05 M real‑world videos. It generates shot‑level cinematic plans through video‑grounded reverse construction and uses Semantic‑Consistency GRPO (SC‑GRPO) to maintain user intent across shots and time. In evaluations, WanPE improves human preference over raw prompts by up to 50.86 points for 30‑second videos and outperforms commercial offerings for shorter durations.
By Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong
arXiv:2606. 30145v1 Announce Type: new Abstract: Natural face-to-face conversation requires real-time speech generation together with synchronized facial motion.
By Habin Lim, Jae-Ho Lee, Hah Min Lew, Ji-Su Kang, Gyeong-Moon Park
arXiv:2609.13830v1 Announce Type: new
Abstract: We present DiVA, a deeply interactive digital life simulator pioneering a new paradigm for long-term, open-ended interactive experiences within digital...
By Cheng Chen, Hao Ouyang, Qiuyu Wang, Ka Leong Cheng, Wen Wang, Yihao Meng, Hanlin Wang, Yixuan Li, Jiacheng Wei, Zhenshan Tan, Yanhong Zeng, Yujun Shen, Guosheng Lin, Fayao Liu
StreamAV-Bench is the first comprehensive benchmark designed for streaming audio‑video generation, addressing the limitations of existing benchmarks that focus on completed sequences. It introduces a unified evaluation framework with a progressive track for instruction adherence and long‑horizon stability, and an interactive track for responsive interaction and state retention. The benchmark includes 32 fine‑grained, expert‑verified evaluation cases and evaluates 13 representative systems, revealing temporal drift in progressive generation and responsiveness bottlenecks in interactive control.
By Kaiqi Liu, Haoxuan Zeng, Jingqi Liu, Jiacong Fang, Ziqi Cai, Yunyao Mao, Henglin Liu, Yu Sheng, Shuchen Weng, Boxin Shi
The paper introduces a real‑time framework for generating co‑speech gestures for digital humans, coupling a streaming speech response module with a causal multimodal autoregressive gesture generator that uses only current speech and motion history. It also presents an offline data synthesis pipeline for virtual companion dialogues and a self‑evolving training loop that incorporates user feedback to continually adapt the model. Experiments show the system achieves a better latency‑quality trade‑off, stronger speech‑motion synchronization, and higher user preference than existing baselines.
By Wentao Jiang, Youchen Xie, Haidi Fan, Yajing Chen, Xin Wang, Ye Shi, Jingya Wang