arXiv Machine Learning

Contrastive Diffusion Alignment: Learning Structured Latents for Controllable Generation

arXiv:2510. 14190v3 Announce Type: replace Abstract: Diffusion models excel at generation, but their latent spaces are high dimensional and not explicitly organized for interpretation or control.

arXiv Computer Vision
Sep 23

Latent Dataset Distillation for Human Motion Prediction

The paper introduces a latent dataset distillation framework for human motion prediction, addressing the limitations of traditional gradient matching by incorporating a learned motion prior. Motions are compressed using a residual‑quantized variational autoencoder, and distillation updates only a latent bank while keeping the decoder frozen, ensuring synthetic motions remain plausible. Experiments on Human3.6M, CMU, and 3DPW datasets demonstrate that this method outperforms direct gradient matching in most settings and yields more realistic synthetic motions.

By Ge Tian, Guang Li, Takahiro Ogawa, Miki Haseyama
arXiv Computer Vision
Sep 16

SlotDiT: Object-Centric Representations for Diffusion Transformers

SlotDiT introduces a text-guided Diffusion Transformer that operates in a slot-based latent space, decomposing scenes into object-centric slots and autoregressively denoising future slot trajectories to predict scene dynamics. The model is conditioned on a reference image and a language instruction, enabling it to generate video content that reflects both visual context and textual guidance. Experiments comparing slot-based representations to VAE-based and semantics-aligned alternatives show that SlotDiT achieves competitive video generation quality while improving task-completion rates across four robotic datasets and offering a more computationally efficient latent representation.

By Gjergj Plepi, Sven Behnke
arXiv AI
Aug 28

The Principles of Diffusion Models

The book "The Principles of Diffusion Models" outlines the foundational concepts behind diffusion models, tracing their evolution from a forward process that corrupts data into noise to a reverse process that reconstructs data. It presents three complementary perspectives—variational, score-based, and flow-based—each describing how a time-dependent velocity field transports a simple prior to the data distribution. The text also covers practical guidance for controllable generation, efficient solvers, and diffusion-inspired flow-map models, providing a mathematically grounded framework for readers with basic deep‑learning knowledge.

By Chieh-Hsin Lai, Yang Song, Dongjun Kim, Yuki Mitsufuji, Stefano Ermon
arXiv AI
Aug 25

Divisive Normalization Shapes Low-Rank Slow Manifolds for Continuous Working Memory

The paper introduces the Recurrent Divisive Normalization Network (RDNN), a minimal model that incorporates divisive normalization—a common neural computation—to stabilize continuous working memory representations. Dynamical systems analysis shows that this biophysical constraint enables the network to converge to robust, high‑fidelity slow manifolds, while gradient dynamics during Backpropagation Through Time reveal an activity‑dependent local scaling that compresses the network’s effective rank into a low‑dimensional subspace. Ablation studies confirm that divisive normalization, rather than subtractive inhibition, is essential for preventing manifold shattering under time‑varying inputs.

By Zhaotian Gu, Jie Su, Weiwei Wang, Chang Liu, Tianyi Qian, Dahui Wang