arXiv Computer Vision

SlotDiT: Object-Centric Representations for Diffusion Transformers

SlotDiT introduces a text-guided Diffusion Transformer that operates in a slot-based latent space, decomposing scenes into object-centric slots and autoregressively denoising future slot trajectories to predict scene dynamics. The model is conditioned on a reference image and a language instruction, enabling it to generate video content that reflects both visual context and textual guidance. Experiments comparing slot-based representations to VAE-based and semantics-aligned alternatives show that SlotDiT achieves competitive video generation quality while improving task-completion rates across four robotic datasets and offering a more computationally efficient latent representation.

arXiv AI
Sep 3

Bernini: Latent Semantic Planning for Video Diffusion

Bernini proposes a unified framework that separates semantic planning and pixel rendering for video generation and editing. An MLLM-based planner predicts target semantics in ViT embedding space, while a DiT-based renderer synthesizes pixels conditioned on this plan, text features, and source VAE features for editing. The approach introduces Segment-Aware 3D Rotary Positional Embedding and chain-of-thought reasoning, achieving state‑of‑the‑art performance on diverse video benchmarks.

By Bernini Team, Chenchen Liu, Junyi Chen, Lei Li, Lu Chi, Mingzhen Sun, Zhuoying Li, Yi Fu, Ruoyu Guo, Yiheng Wu, Ge Bai, Zehuan Yuan
arXiv AI
Sep 17

From Alignment to Synthesis: Contrastive Volumetric Grounding for Text-to-CT Generation

The paper introduces a 3D-CLIP encoder trained with structured hard negatives to improve vision‑language alignment for text‑to‑CT generation. This encoder drives a latent diffusion model that operates directly in 3D latent space, eliminating spatial artifacts from super‑resolution pipelines. Experiments on the CT‑RATE dataset show state‑of‑the‑art image fidelity and factual correctness across 18 pathological conditions, with lower inference time and GPU memory usage than competing methods.

By Daniele Molino, Camillo Maria Caruso, Filippo Ruffini, Paolo Soda, Valerio Guarrasi
arXiv Computer Vision
Sep 24

AWM-VLA: AlignedWorld Modeling for Efficient and Explainable Vision-Language-Action Policies

AWM‑VLA introduces a unified framework that embeds aligned world modeling directly into a diffusion‑transformer vision‑language‑action policy. By adding learnable future tokens aligned with vision‑language embeddings of future observations, the policy can anticipate long‑term consequences while generating actions. The method extends this with an object‑centric alignment objective and a principled weighting scheme, achieving up to 21% higher success rates on RoboCasa and humanoid tabletop benchmarks and producing object‑centric rationales preferred by human raters in 83% of cases.

By An Lanji, Dawei Liu, Jin Li, Haoran Xu, Mei Chen, Yu Tian
arXiv AI
Sep 10

VoT: Vision-of-Thought for Unified Multimodal Representation Alignment

The paper introduces Vision-of-Thought (VoT), a framework that inserts a discrete visual-thinking layer between vision‑language models (VLMs) and diffusion transformers (DiTs). Instead of using VLMs solely as text encoders, they act as multimodal planners that generate VoT tokens—high‑level visual plans such as objects and layouts—before pixel rendering. A specialized VoT tokenizer is trained with a closed‑loop objective combining VLM alignment, feature reconstruction, and vector‑quantization losses, ensuring the tokens are semantically readable by the VLM while preserving necessary visual information. Experimental results show that VoT improves semantic alignment and offers a structured, interpretable interface for controllable generation.

By Jingxiang Sun, Chao Liao, Zhengxiong Luo, Chaorui Deng, Chen-lin Zhang, Junke Wang, Ceyuan Yang, Haoqi Fan, Weilin Huang