arXiv:2609.08032v1 Announce Type: cross
Abstract: We introduce FlexMoGen, a novel framework for flexible human motion synthesis conditioned on both natural language descriptions and motion style refe...
By Kai Weixian Lan, Bodie Criswell, Briana Fedkiw, Zhan Zhang, Joseph Teran, Daniel Holden
arXiv:2609.23817v1 Announce Type: new
Abstract: We present VISTA, a two-stage framework for generating stylized 3D human motion by fusing structural content from text prompts with expressive style fr...
By Monseej Purkayastha, Anindita Ghosh, Philipp Slusallek
arXiv:2609.25775v1 Announce Type: new
Abstract: Recent advances in generative video models have enabled the synthesis of visually realistic content, posing significant challenges to synthetic video d...
By Huangsen Cao, Hongkang chu, Siyao Yu, Xin Ding, Jianfeng Dong, Yongwei Wang
arXiv:2606. 19676v1 Announce Type: cross Abstract: Diffusion models have achieved remarkable success in image and video generation and editing.
By Haengbok Chung
arXiv:2510.24904v2 Announce Type: replace
Abstract: Although recent video generative models are getting more capable of following external camera controls, imposed by either text descriptions or came...
By Qiucheng Wu, Handong Zhao, Zhixin Shu, Jing Shi, Yang Zhang, Shiyu Chang
arXiv:2605. 23045v2 Announce Type: replace-cross Abstract: Video representation learning has seen tremendous progress in recent years.
By Mantas Skackauskas, Xinyue Hao, Laura Sevilla-Lara
arXiv:2608.20770v1 Announce Type: new
Abstract: Modern AI video generation models can produce videos with high visual fidelity and seemingly smooth temporal transitions. However, visual realism does...
By Haojin He, Hao Tan, Zichang Tan, Ajian Liu, Jun Wan
ContextAnyone is a context‑aware diffusion framework that treats a reference image as an explicitly preserved appearance anchor rather than a simple conditioning signal. By jointly reconstructing the reference image and generating the target video within a shared diffusion transformer, it provides direct supervision for maintaining identity and fine‑grained appearance throughout denoising. The method introduces asymmetric information flow and Gap‑RoPE positional representations to keep the reference stable while allowing selective access by video tokens, and demonstrates improved identity and appearance consistency on an OpenVid‑HD benchmark.
By Ziyang Mai, Yu-Wing Tai
arXiv:2609.22267v1 Announce Type: new
Abstract: Camera motion often reflects directorial intent and requires professional equipment, making it a high value form of intellectual property. However, gen...
By Chengguo Zhang, Ping Ping
arXiv:2610.00812v1 Announce Type: cross
Abstract: Video generation has rapidly progressed from short, low-quality clips to high-resolution, long-duration sequences with complex spatiotemporal dynamic...
By Chaoyu Li, Xiaoyi Gu, Yogesh Kulkarni, Eun Woo Im, Mohammadmahdi Honarmand, Zeyu Wang, Juntong Song, Fei Du, Xilin Jiang, Kexin Zheng, Tianzhi Li, Fei Tao, Pooyan Fazli
arXiv:2607. 22830v2 Announce Type: replace Abstract: In visual storytelling, human performances are central to creative intent and narrative meaning.
By Yuancheng Xu, Mingming He, Pablo Salamanca, Li Ma, Yash Kant, Emmett Steven, Paul Debevec, Ning Yu
The paper compares diffusion and rectified flow objectives within the MotionGPT3 framework for text-driven motion generation. Experiments on HumanML3D show that rectified flow converges faster, achieves strong test performance earlier, and matches or exceeds diffusion quality while requiring fewer sampling steps. The study isolates the generative objective’s impact, demonstrating that rectified flow’s benefits transfer to continuous-latent motion generation.
By Jaymin Bhan, JiHong Jeon, SangYeop Jeong