WanPE is a 397‑B parameter prompt‑enhancement model that learns director‑level cinematic planning from 1.05 M real‑world videos. It generates shot‑level cinematic plans through video‑grounded reverse construction and uses Semantic‑Consistency GRPO (SC‑GRPO) to maintain user intent across shots and time. In evaluations, WanPE improves human preference over raw prompts by up to 50.86 points for 30‑second videos and outperforms commercial offerings for shorter durations.
By Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong
arXiv:2607. 06481v1 Announce Type: cross Abstract: We present PACR-Video, a parameter-efficient framework for multi-shot long video extrapolation that preserves recurring entities, scene structure, visual style, and causal progression without full generator fine-tuning.
By Anna C\'ordoba, Adam Puente Tercero, Nerea Angulo Hijo, Mar Linares Tercero, Julia Barrientos, Ainhoa Miranda, Jes\'us Olivera
arXiv:2609.36496v1 Announce Type: new
Abstract: Most video editing methods focus on changing the appearance of the source video, while offering limited control over its dynamics. We introduce Reimagi...
By Yu Yuan, Yawen Lu, Guoxian Song, Kevin Duarte, Ratheesh Kalarot, Di Chang, Xijun Wang, Stanley H. Chan
arXiv:2609.37495v1 Announce Type: new
Abstract: Human motion generation plays an important role in applications such as character animation, virtual environments, and embodied interaction. While exis...
By Yun Chen, Munchurl Kim, Jeonghyeok Do
LIFT is a unified image‑to‑video generation framework that adds Layout‑In‑Future control, letting users specify what should appear and where in a future view. It addresses the limitation of existing camera controls and text prompts by using the last‑frame layout as an explicit signal for the desired future scene, especially under large viewpoint changes. To handle sparse layout guidance, LIFT employs on‑policy self‑distillation to transfer knowledge from a dense‑layout teacher to a last‑frame‑layout student, and introduces the LIFT‑Vista dataset with large viewpoint changes and consistent layout annotations. Experiments demonstrate that LIFT improves video quality, future‑layout controllability, and camera controllability compared to other methods.
By Shengxiang Ji, Boyang Wang, Haiyang Xu, Bingnan Li, Yucheng Mao, Zeyuan Chen, Xiaojun Shan, Xiang Zhang, Gang Hua, Jianwen Xie, Zezhou Cheng, Zhuowen Tu
Text-conditioned image-to-video (I2V) generation has advanced rapidly, yet generating videos with multiple subjects remains challenging. A model must simultaneously preserve the appearance of each sub...
MoVT is a new framework for text‑to‑motion generation that uses a cross‑modal augmented motion tokenizer to project 3D motion tokens into 2D, enriching the motion codebook with real‑world video patterns. The enriched tokens are mapped back to 3D, creating aligned 3D and 2D codebooks that better capture intricate motions. These codebooks feed a generative masked transformer, which predicts masked motion tokens in a modality‑agnostic way, allowing text‑index pairs from the 2D codebook and annotated videos to further improve generation quality. Empirical tests show MoVT outperforms previous state‑of‑the‑art methods on several key metrics.
By Beibei Jing, Tianle Guo, Youjia Zhang, Zikai Song, Yawei Luo, Junqing Yu, Tao Guan, Wei Yang
arXiv:2608.22819v1 Announce Type: new
Abstract: Text-conditioned image-to-video (I2V) generation has advanced rapidly, yet generating videos with multiple subjects remains challenging. A model must s...
By Yanliang Qi, Kexi Chen, Muchao Ye, Haomiao Ni
EditaLive! is a new real‑time framework for character video editing in live streaming, built on a pretrained image animation model (Wan‑Animate) that separates appearance from motion. It uses the CharEdit‑50K dataset for reference‑frame editing and video reconstruction, and adapts the model from offline bidirectional to causal streaming generation. A self‑rollout distillation strategy compresses the model into a two‑step sampler, employing fixed RoPE, alignment forcing, and first‑frame preserved sparse attention to reduce appearance drift and achieve low‑latency inference while preserving facial expressions.
By Zhiyuan Li, Chi-Man Pun, Peng-Tao Jiang, Bo Li, Xiaodong Cun
EditStream is a unified DiT‑based framework that supports a wide range of interactive video tasks—Text‑to‑Video, Image‑to‑Video, Video‑to‑Video, Editing Propagation, Reference‑guided Video Editing, and Camera Pose Change—within a single system. It achieves fast, few‑step autoregressive generation by applying a two‑stage distillation process that combines Velocity Moment Matching with autoregressive unrolling, thereby preserving motion quality and temporal stability. The approach aims to make high‑quality diffusion‑based video models practical for real‑time creative workflows.
By Yuqian Zhou, Zhenghong Zhou, Zongze Wu, Cameron Smith, Richard Zhang, Jiebo Luo, Eli Shechtman, Zhe Lin
Memory-V2V is a memory‑augmented video‑to‑video diffusion framework designed to improve cross‑turn consistency in multi‑turn video editing. It stores previous outputs in an external memory, retrieves relevant edits, and incorporates them via relevance‑aware tokenization and adaptive compression, allowing scalable conditioning without linear computational growth. Experiments on iterative video novel view synthesis and text‑guided long video editing show that Memory‑V2V enhances consistency while preserving visual quality and outperforming strong baselines with modest overhead.
By Dohun Lee, Chun-Hao Paul Huang, Xuelin Chen, Jong Chul Ye, Duygu Ceylan, Hyeonho Jeong
arXiv:2609.36598v1 Announce Type: new
Abstract: A video can exhibit convincing motion and photorealism yet fail immediately when visual text collapses. Unlike generic scene content, visual text is un...
By Ziying Zhang, Litao Li, Junchao Liao, Tianyi Zeng, Siyu Zhu, Long Qin, Zhenghao Zhang