ReaDiT Guidance is a lightweight framework that controls Diffusion Transformer (DiT) models by leveraging internal feature representations from a single DiT block. It steers image and video generation toward spatial targets such as depth, pose, or edge maps, and extends naturally to video generation for camera and motion control. Experiments show competitive or improved results compared to existing feature‑based and adapter‑based methods while using fewer parameters.
By Jay Mahajan, Chang Liu, Rauf Makharov, Viraj Shah, Alexander Schwing, Svetlana Lazebnik
BulletTime presents a 4D‑controllable video diffusion framework that separates scene dynamics from camera pose, allowing precise manipulation of both temporal and spatial aspects of generated videos. The model conditions on continuous world‑time sequences and camera trajectories, integrating them via a 4D positional encoding in the attention layer and adaptive normalizations for feature modulation. A specially curated dataset with independently parameterized temporal and camera variations is used for training, and the resulting system demonstrates robust real‑world 4D control while maintaining high generation quality and surpassing prior methods in controllability.
By Yiming Wang, Qihang Zhang, Shengqu Cai, Tong Wu, Jan Ackermann, Zhengfei Kuang, Yang Zheng, Frano Raji\v{c}, Siyu Tang, Gordon Wetzstein
arXiv:2606. 06853v1 Announce Type: cross Abstract: The new era has witnessed a remarkable capability to extend Vision-Language Models (VLMs) for tackling tasks of video understanding.
By Yifan Xu, Chao Zhang, Ruifei Ma, Fei Gao, Zhifei Yang, Jiaxing Qi, Zhipeng Chen
arXiv:2606. 13432v1 Announce Type: cross Abstract: Cloning camera motion from reference videos is an important task in video generation, as videos provide intuitive and precise control.
By Jiwen Liu, Shujuan Li, Zhixue Fang, Xiaohan Li, Yan Zhou, Zijie Meng, Zhimin Zhang, Yawen Luo, Guoxin Zhang, Yu-Shen Liu, Pengfei Wan
The paper tackles two main issues in multi‑subject video generation—uncontrollable fidelity strength and semantic drift—by exploiting intrinsic attention patterns in Diffusion Transformers. It introduces an Intrinsic Spatial Grounding Map (ISGM) that accurately locates reference subjects and a Dual‑phase Intrinsic Attention Leveraging (DIAL) framework that uses ISGM during both training and inference. DIAL guides attention in low‑noise stages for precise fidelity control and builds preference pairs in high‑noise stages for reinforcement learning, resulting in superior identity consistency and controllable fidelity on the OpenS2V‑Eval benchmark.
By Niange Yu, Ye Tian, Biaolong Chen, Miao Lu, Aixi Zhang, Hao Jiang, Yunhai Tong, Pipei Huang
ContextAnyone is a context‑aware diffusion framework that treats a reference image as an explicitly preserved appearance anchor rather than a simple conditioning signal. By jointly reconstructing the reference image and generating the target video within a shared diffusion transformer, it provides direct supervision for maintaining identity and fine‑grained appearance throughout denoising. The method introduces asymmetric information flow and Gap‑RoPE positional representations to keep the reference stable while allowing selective access by video tokens, and demonstrates improved identity and appearance consistency on an OpenVid‑HD benchmark.
By Ziyang Mai, Yu-Wing Tai
arXiv:2602.24289v2 Announce Type: replace
Abstract: Scaling video generation from seconds to minutes faces a critical bottleneck: while short-video data is abundant and high-fidelity, coherent long-f...
By Shengqu Cai, Weili Nie, Chao Liu, Julius Berner, Lvmin Zhang, Nanye Ma, Hansheng Chen, Maneesh Agrawala, Leonidas Guibas, Gordon Wetzstein, Arash Vahdat
arXiv:2510.24904v2 Announce Type: replace
Abstract: Although recent video generative models are getting more capable of following external camera controls, imposed by either text descriptions or came...
By Qiucheng Wu, Handong Zhao, Zhixin Shu, Jing Shi, Yang Zhang, Shiyu Chang
arXiv:2606. 10183v1 Announce Type: cross Abstract: Modern Diffusion Transformers for video generation provide limited control over the progression of time and the editing of temporal dynamics.
By Konstantin Kuklev, Viacheslav Vasilev, Alexander Kunitsyn, Andrei Ivaniuta, Denis Dimitrov
The paper tackles two main issues in multi-subject video generation—uncontrollable fidelity strength and semantic drift—by studying Diffusion Transformers (DiTs). It discovers that certain attention blocks naturally create an Intrinsic Spatial Grounding Map (ISGM) that accurately locates reference subjects. Leveraging this insight, the authors introduce Dual-phase Intrinsic Attention Leveraging (DIAL), which uses ISGM during low-noise stages to control fidelity strength without retraining and during high-noise stages to generate preference pairs for reinforcement learning, thereby anchoring attention and reducing semantic drift. Experiments on the OpenS2V-Eval benchmark show that DIAL outperforms baseline models, improving identity consistency and enabling controllable fidelity strength.
arXiv:2506.01004v3 Announce Type: replace-cross
Abstract: Unlike traditional video editing or inpainting, video semantic mixing fuses a reference concept with a moving target entity to produce a hybr...
By Tong Zhang, Victor Escorcia, Juan C Leon Alcazar, Bernard Ghanem
Synthesizing a novel-view video from a monocular reference video along a target camera trajectory requires both geometric consistency and motion fidelity with respect to the reference video. Existing methods based on explicit 3D representations are limited by the accuracy of off-the-shelf reconstruction modules, which often produce inaccurate geometry for dynamic objects in monocular videos.