arXiv:2608.20770v1 Announce Type: new
Abstract: Modern AI video generation models can produce videos with high visual fidelity and seemingly smooth temporal transitions. However, visual realism does...
By Haojin He, Hao Tan, Zichang Tan, Ajian Liu, Jun Wan
arXiv:2604.09057v3 Announce Type: replace
Abstract: Audio-video (AV) generation has recently made strong progress in perceptual quality and multimodal coherence, yet generating content with plausible...
By Junchao Liao, Zhenghao Zhang, Xiangyu Meng, Litao Li, Ziying Zhang, Siyu Zhu, Long Qin, Weizhi Wang
arXiv:2609.37495v1 Announce Type: new
Abstract: Human motion generation plays an important role in applications such as character animation, virtual environments, and embodied interaction. While exis...
By Yun Chen, Munchurl Kim, Jeonghyeok Do
We propose Symmetric Nonlinear Motion-guided Generative Video Frame Interpolation (SNM-VFI), a training-free framework for motion-controllable generative video frame interpolation with pre-trained optical flow and video diffusion models. Unlike conventional diffusion-based VFI methods that synthesize intermediate frames from random noise, SNM-VFI guides the generative process with correspondence-aware frames produced by a symmetric nonlinear motion model.
arXiv:2609.25775v1 Announce Type: new
Abstract: Recent advances in generative video models have enabled the synthesis of visually realistic content, posing significant challenges to synthetic video d...
By Huangsen Cao, Hongkang chu, Siyao Yu, Xin Ding, Jianfeng Dong, Yongwei Wang
arXiv:2609.23817v1 Announce Type: new
Abstract: We present VISTA, a two-stage framework for generating stylized 3D human motion by fusing structural content from text prompts with expressive style fr...
By Monseej Purkayastha, Anindita Ghosh, Philipp Slusallek
The paper compares diffusion and rectified flow objectives within the MotionGPT3 framework for text-driven motion generation. Experiments on HumanML3D show that rectified flow converges faster, achieves strong test performance earlier, and matches or exceeds diffusion quality while requiring fewer sampling steps. The study isolates the generative objective’s impact, demonstrating that rectified flow’s benefits transfer to continuous-latent motion generation.
By Jaymin Bhan, JiHong Jeon, SangYeop Jeong
CameraEditor is a new framework that transforms camera-controlled image editing into a temporal sequence prediction problem. By using video diffusion models, it incorporates a geometric perception module and dynamic reference routing to create precise visual references through dynamic panorama cropping. The method also inserts intermediate transition frames to handle large perspective shifts, maintaining content identity and spatial coherence, and is evaluated on a dataset of 5,760 instances with a benchmark of 462 test cases, achieving state‑of‑the‑art performance.
By Xin Shen, Chengyou Jia, Keshuo Xing, Zifeng Zhu, Changliang Xia, Bowen Ping, Zhuohang Dang, Hangwei Qian, Minnan Luo
arXiv:2608.22861v1 Announce Type: new
Abstract: State Space Models (SSMs) have surfaced as a promising architecture in Video Frame Interpolation (VFI), as they can capture long-range dependencies wit...
By Jaehyun Park, Nam Ik Cho
arXiv:2609.36496v1 Announce Type: new
Abstract: Most video editing methods focus on changing the appearance of the source video, while offering limited control over its dynamics. We introduce Reimagi...
By Yu Yuan, Yawen Lu, Guoxian Song, Kevin Duarte, Ratheesh Kalarot, Di Chang, Xijun Wang, Stanley H. Chan
Streaming autoregressive diffusion models enable real-time, long-horizon video generation, but their training objectives optimize local frame prediction rather than the geometry and dynamics of a coherent world: long rollouts accumulate geometric drift and degrade into static or unnatural motion. Recent bidirectional approaches address this problem using rewards signals built upon 3D Gaussian-Splatting reconstruction.
arXiv:2603. 22282v2 Announce Type: replace-cross Abstract: We present UniMotion, to our knowledge the first unified framework for simultaneous understanding and generation of human motion, natural language, and RGB images within a single architecture.
By Ziyi Wang, Xinshun Wang, Shuang Chen, Yang Cong, Mengyuan Liu