arXiv AI

Making Time Editable in Video Diffusion Transformers

arXiv:2606. 10183v1 Announce Type: cross Abstract: Modern Diffusion Transformers for video generation provide limited control over the progression of time and the editing of temporal dynamics.

arXiv Computer Vision
Aug 25

EditStream: A Unified Autoregressive Framework for Interactive Video Generation and Editing

EditStream is a unified DiT‑based framework that supports a wide range of interactive video tasks—Text‑to‑Video, Image‑to‑Video, Video‑to‑Video, Editing Propagation, Reference‑guided Video Editing, and Camera Pose Change—within a single system. It achieves fast, few‑step autoregressive generation by applying a two‑stage distillation process that combines Velocity Moment Matching with autoregressive unrolling, thereby preserving motion quality and temporal stability. The approach aims to make high‑quality diffusion‑based video models practical for real‑time creative workflows.

By Yuqian Zhou, Zhenghong Zhou, Zongze Wu, Cameron Smith, Richard Zhang, Jiebo Luo, Eli Shechtman, Zhe Lin
arXiv AI
1d ago

VIDiff: Translating Videos via Multi-Modal Instructions with Diffusion Models

VIDiff is a unified foundation model that uses diffusion techniques to perform a broad range of video tasks, including both understanding tasks like language‑guided video object segmentation and generative tasks such as video editing and enhancement. Unlike prior methods that focus on short clips and require time‑consuming tuning, VIDiff can edit and translate videos within seconds based on user instructions and employs an iterative auto‑regressive approach to maintain consistency in long‑form videos. The authors demonstrate convincing generative results across diverse input videos and written instructions, supported by qualitative and quantitative evidence.

By Zhen Xing, Shuyuan Tu, Qi Dai, Zihao Zhang, Hui Zhang, Han Hu, Zuxuan Wu, Yu-Gang Jiang
Hugging Face Trending Papers
Aug 4

JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion

Real-time video editing requires low-latency causal generation with bounded computational resources while preserving source fidelity and long-term temporal consistency. We present JoyAI-Video-Edit, a 16B-parameter autoregressive diffusion framework for real-time, open-ended video editing without access to future frames or a predefined video duration.

arXiv Computer Vision
Aug 25

BulletTime: Decoupled Control of Time and Camera Pose for Video Generation

BulletTime presents a 4D‑controllable video diffusion framework that separates scene dynamics from camera pose, allowing precise manipulation of both temporal and spatial aspects of generated videos. The model conditions on continuous world‑time sequences and camera trajectories, integrating them via a 4D positional encoding in the attention layer and adaptive normalizations for feature modulation. A specially curated dataset with independently parameterized temporal and camera variations is used for training, and the resulting system demonstrates robust real‑world 4D control while maintaining high generation quality and surpassing prior methods in controllability.

By Yiming Wang, Qihang Zhang, Shengqu Cai, Tong Wu, Jan Ackermann, Zhengfei Kuang, Yang Zheng, Frano Raji\v{c}, Siyu Tang, Gordon Wetzstein
arXiv Computer Vision
Sep 7

ReaDiT Guidance: Control for Image and Video Generation using Diffusion Transformer Features

ReaDiT Guidance is a lightweight framework that controls Diffusion Transformer (DiT) models by leveraging internal feature representations from a single DiT block. It steers image and video generation toward spatial targets such as depth, pose, or edge maps, and extends naturally to video generation for camera and motion control. Experiments show competitive or improved results compared to existing feature‑based and adapter‑based methods while using fewer parameters.

By Jay Mahajan, Chang Liu, Rauf Makharov, Viraj Shah, Alexander Schwing, Svetlana Lazebnik
arXiv Computer Vision
6d ago

Reimagine Video Dynamics

arXiv:2609.36496v1 Announce Type: new Abstract: Most video editing methods focus on changing the appearance of the source video, while offering limited control over its dynamics. We introduce Reimagi...

By Yu Yuan, Yawen Lu, Guoxian Song, Kevin Duarte, Ratheesh Kalarot, Di Chang, Xijun Wang, Stanley H. Chan
arXiv Machine Learning
Aug 27

Memory-V2V: Memory-Augmented Video-to-Video Diffusion for Consistent Multi-Turn Editing

Memory-V2V is a memory‑augmented video‑to‑video diffusion framework designed to improve cross‑turn consistency in multi‑turn video editing. It stores previous outputs in an external memory, retrieves relevant edits, and incorporates them via relevance‑aware tokenization and adaptive compression, allowing scalable conditioning without linear computational growth. Experiments on iterative video novel view synthesis and text‑guided long video editing show that Memory‑V2V enhances consistency while preserving visual quality and outperforming strong baselines with modest overhead.

By Dohun Lee, Chun-Hao Paul Huang, Xuelin Chen, Jong Chul Ye, Duygu Ceylan, Hyeonho Jeong