arXiv Computer Vision

Diffusion Trajectory Modeling for Semantic Correspondence

Diffusion Trajectory Modeling (DTM) treats the evolving feature maps of diffusion models as temporally structured trajectories rather than static snapshots. By interpreting each spatial patch’s progression across multiple timesteps as a trajectory, DTM captures semantic correspondence cues that prior methods miss. Experiments on SPair-71k, SPair-U, and AP-10K demonstrate that DTM achieves strong performance, highlighting the semantic value embedded in the diffusion process’s temporal axis.

arXiv Computer Vision
Aug 26

Representation Learning in Diffusion and Flow-based Model: An Application Aspect

The article surveys how diffusion and flow-based generative models learn rich visual representations and how these representations can be used to improve generation and other perception tasks. It introduces a three-tier framework that categorizes work into improving generative quality via representation learning, extracting representations for perception, and developing unified applications. The survey covers downstream tasks such as image classification, dense prediction, instance-level perception, and annotation-scarce scenarios, offering a taxonomy and highlighting future research directions.

By Yanchen Xu, Sida Huang, Zhenyu Gu, Ruishu Zhu, Yilan Gao, Hongyuan Zhang
arXiv AI
Jul 8

Few Channels Draw The Whole Picture: Revealing Massive Activations in Diffusion Transformers

arXiv:2605. 13974v2 Announce Type: replace-cross Abstract: Diffusion Transformers (DiTs) and related flow-based architectures are now among the strongest text-to-image generators, yet the internal mechanisms through which prompts shape image semantics remain poorly understood.

By Evelyn Turri, Davide Bucciarelli, Sara Sarto, Lorenzo Baraldi, Marcella Cornia
arXiv Computer Vision
Sep 22

AlignMorph: Tuning-Free Diffusion Image Morphing via Explicit Semantic Transport

AlignMorph is a tuning‑free diffusion framework for image morphing that separates geometric alignment from generative denoising. It uses Global Semantic Transport—entropic optimal transport and reliability‑aware latent warping—to achieve diffusion‑compatible semantic alignment, and Coordinate‑Aligned Generation—symmetric bi‑phase attention handoff—to preserve spatial coordinates during denoising. The method eliminates ghosting and delivers superior structural coherence and temporal smoothness on morphing benchmarks without any per‑pair optimization.

By Wuyi Liu, Xu Han, Yuren Chen, Yige Mao, Zishuo Peng, Xianzhi Li
arXiv Computer Vision
Sep 16

SlotDiT: Object-Centric Representations for Diffusion Transformers

SlotDiT introduces a text-guided Diffusion Transformer that operates in a slot-based latent space, decomposing scenes into object-centric slots and autoregressively denoising future slot trajectories to predict scene dynamics. The model is conditioned on a reference image and a language instruction, enabling it to generate video content that reflects both visual context and textual guidance. Experiments comparing slot-based representations to VAE-based and semantics-aligned alternatives show that SlotDiT achieves competitive video generation quality while improving task-completion rates across four robotic datasets and offering a more computationally efficient latent representation.

By Gjergj Plepi, Sven Behnke
arXiv AI
Sep 10

VoT: Vision-of-Thought for Unified Multimodal Representation Alignment

The paper introduces Vision-of-Thought (VoT), a framework that inserts a discrete visual-thinking layer between vision‑language models (VLMs) and diffusion transformers (DiTs). Instead of using VLMs solely as text encoders, they act as multimodal planners that generate VoT tokens—high‑level visual plans such as objects and layouts—before pixel rendering. A specialized VoT tokenizer is trained with a closed‑loop objective combining VLM alignment, feature reconstruction, and vector‑quantization losses, ensuring the tokens are semantically readable by the VLM while preserving necessary visual information. Experimental results show that VoT improves semantic alignment and offers a structured, interpretable interface for controllable generation.

By Jingxiang Sun, Chao Liao, Zhengxiong Luo, Chaorui Deng, Chen-lin Zhang, Junke Wang, Ceyuan Yang, Haoqi Fan, Weilin Huang