arXiv AI
Sep 3

Bernini: Latent Semantic Planning for Video Diffusion

Bernini proposes a unified framework that separates semantic planning and pixel rendering for video generation and editing. An MLLM-based planner predicts target semantics in ViT embedding space, while a DiT-based renderer synthesizes pixels conditioned on this plan, text features, and source VAE features for editing. The approach introduces Segment-Aware 3D Rotary Positional Embedding and chain-of-thought reasoning, achieving state‑of‑the‑art performance on diverse video benchmarks.

By Bernini Team, Chenchen Liu, Junyi Chen, Lei Li, Lu Chi, Mingzhen Sun, Zhuoying Li, Yi Fu, Ruoyu Guo, Yiheng Wu, Ge Bai, Zehuan Yuan
arXiv AI
5d ago

VIDiff: Translating Videos via Multi-Modal Instructions with Diffusion Models

VIDiff is a unified foundation model that uses diffusion techniques to perform a broad range of video tasks, including both understanding tasks like language‑guided video object segmentation and generative tasks such as video editing and enhancement. Unlike prior methods that focus on short clips and require time‑consuming tuning, VIDiff can edit and translate videos within seconds based on user instructions and employs an iterative auto‑regressive approach to maintain consistency in long‑form videos. The authors demonstrate convincing generative results across diverse input videos and written instructions, supported by qualitative and quantitative evidence.

By Zhen Xing, Shuyuan Tu, Qi Dai, Zihao Zhang, Hui Zhang, Han Hu, Zuxuan Wu, Yu-Gang Jiang
arXiv AI
Sep 30

PreviewDiff: Multimodal Critic-Guided Search over Diffusion Latents

PreviewDiff is a test‑time search method that uses multimodal critics to guide diffusion model sampling. By decoding partial previews at selected denoising checkpoints, scoring them with a multimodal judge, and branching over semantic prompt edits, it allows the generation process to be edited and rerouted before completion. The approach consistently outperforms budget‑matched Best‑of‑N sampling and scalar‑search baselines on image and video benchmarks, with early interventions and wider search yielding the biggest gains.

By Vighnesh Subramaniam, Boris Katz, Brian Cheung, Chun-Liang Li, Tomas Pfister, Yale Song