arXiv Computer Vision

ReaDiT Guidance: Control for Image and Video Generation using Diffusion Transformer Features

ReaDiT Guidance is a lightweight framework that controls Diffusion Transformer (DiT) models by leveraging internal feature representations from a single DiT block. It steers image and video generation toward spatial targets such as depth, pose, or edge maps, and extends naturally to video generation for camera and motion control. Experiments show competitive or improved results compared to existing feature‑based and adapter‑based methods while using fewer parameters.

arXiv AI
Sep 16

Efficient Text-to-Image Generation: An Adaptive Step Schedule Controller for Diffusion Models

The paper introduces an adaptive step schedule controller for text‑to‑image diffusion models, allowing the number of denoising steps to vary based on the complexity of the input prompt. By mixing step schedules of different sizes and monitoring error discrepancies at each timestep, the method switches schedules to maintain image quality while reducing inference time. Experiments on COCO and DiffusionDB demonstrate that this approach achieves faster generation without sacrificing visual fidelity.

By Kuluhan Binici, Cihan Acar, Shivam Aggarwal, Siying Liu, Tulika Mitra
arXiv Computer Vision
Aug 25

BulletTime: Decoupled Control of Time and Camera Pose for Video Generation

BulletTime presents a 4D‑controllable video diffusion framework that separates scene dynamics from camera pose, allowing precise manipulation of both temporal and spatial aspects of generated videos. The model conditions on continuous world‑time sequences and camera trajectories, integrating them via a 4D positional encoding in the attention layer and adaptive normalizations for feature modulation. A specially curated dataset with independently parameterized temporal and camera variations is used for training, and the resulting system demonstrates robust real‑world 4D control while maintaining high generation quality and surpassing prior methods in controllability.

By Yiming Wang, Qihang Zhang, Shengqu Cai, Tong Wu, Jan Ackermann, Zhengfei Kuang, Yang Zheng, Frano Raji\v{c}, Siyu Tang, Gordon Wetzstein
Hugging Face Trending Papers
Aug 13

Spatially-Grounded Text-to-Video Generation via Inference-Time Gradient-Free Optimization

Diffusion Transformer Text-to-Video models have achieved remarkable synthesis quality, yet fine-grained spatial controllability remains a significant challenge. While existing training-free methods produce solid overall results in spatially grounded generation, \ie, placing a specific object in a designated location, they rely on gradient-based optimization techniques that incur prohibitive computational overhead, a bottleneck amplified in modern large-scale architectures.

arXiv AI
Sep 10

VoT: Vision-of-Thought for Unified Multimodal Representation Alignment

The paper introduces Vision-of-Thought (VoT), a framework that inserts a discrete visual-thinking layer between vision‑language models (VLMs) and diffusion transformers (DiTs). Instead of using VLMs solely as text encoders, they act as multimodal planners that generate VoT tokens—high‑level visual plans such as objects and layouts—before pixel rendering. A specialized VoT tokenizer is trained with a closed‑loop objective combining VLM alignment, feature reconstruction, and vector‑quantization losses, ensuring the tokens are semantically readable by the VLM while preserving necessary visual information. Experimental results show that VoT improves semantic alignment and offers a structured, interpretable interface for controllable generation.

By Jingxiang Sun, Chao Liao, Zhengxiong Luo, Chaorui Deng, Chen-lin Zhang, Junke Wang, Ceyuan Yang, Haoqi Fan, Weilin Huang
arXiv AI
Sep 4

LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes

LLaDA-Image is a unified framework that couples a 6B Diffusion Transformer (DiT) trained from scratch with a frozen vision‑language module based on the LLaDA2.0‑Mini diffusion language model. The approach first builds a strong visual generative prior through image‑only pre‑training and mid‑training, then fine‑tunes with a 220M‑sample generation pipeline that includes 98 real images. The resulting model produces highly photorealistic images that accurately follow fine‑grained editing instructions, and a distilled version, LLaDA‑Image‑Turbo, enables fast inference in 2–4 sampling steps. On Qwen‑Image‑Bench, LLaDA‑Image sets new state‑of‑the‑art scores for open‑source models in both English and Chinese tracks, and the authors release weights, code, and detailed recipes to support further research.

By Chuyan Chen, Haoxing Chen, Kun Chen, Zhenglin Cheng, Long Cui, Ruishan Fang, Zhangxuan Gu, Zhicheng Huang, Zhenzhong Lan, Yuanting Lei, Haoquan Li, Jianguo Li, Rongchuan Li, Sidu Li, Tao Lin, Deyuan Liu, Jiacheng Liu, Lin Liu, Yuxuan Lou, Zhisheng Lu, Yuxin Ma, Shuheng Shen, Peng Sun, Chaoyang Wang, Hongjun Wang, Xiaomei Wang, Yongxin Wang, Chengzhang Wu, Hongru Wu, Jun Xie