arXiv Machine Learning

ORCA: Hunting Compositional Failures in Text-to-Image Diffusion

ORCA (Orthogonal Residual Compositional Alignment) is a method that improves text-to-image diffusion models by aligning the diffusion transformer’s latent space with a low‑rank target derived from a frozen visual encoder. It introduces an auxiliary loss that uses a predictor with an orthogonal basis parameterised by a learned residual between T5 and CLIP embeddings, providing a prompt‑dependent signal for selecting the visual readout subspace. Experiments on three diffusion‑transformer backbones show that ORCA improves FID and GenEval scores, especially on attribute binding, spatial relations, and multi‑object prompts, without adding inference‑time cost.

arXiv Computation and Language
Aug 28

Visual Information-Guided Parallel Decoding for Diffusion Multimodal Large Language Models

Visual Information-Guided Parallel Decoding for Diffusion Multimodal Large Language Models introduces the VIG‑Sampler, a method that prioritizes tokens for decoding based on their attention to image tokens and penalizes redundancy in image‑attention distributions. The approach aims to improve the quality of multimodal generation by selecting more informative tokens during diffusion decoding. Experiments on seven captioning and VQA benchmarks with three open‑source dMLLMs show that VIG‑Sampler outperforms the Info‑Gain Sampler by an average of 19.3 CIDEr points and achieves better COCO Caption results using only half as many decoding steps.

By Insu Lee, Wooje Park, Wonseok Shin, Jinwoo Son, Byonghyo Shim
arXiv Computer Vision
Aug 25

VISTA: Test-Time Compositional Alignment for Visual Autoregressive Generation

VISTA is a gradient‑based test‑time alignment framework designed for next‑scale visual autoregressive (VAR) image generation. It optimizes intermediate representations within the frozen transformer to enforce compositional constraints, without altering model weights or requiring extra training. Experiments on two benchmarks and two model scales show that VISTA improves compositional accuracy by up to 20% on a 2B backbone and 6% on an 8B backbone, while preserving image quality and enabling a smaller model to outperform a larger one.

By Hossein Shahabadi, Niki Sepasian, Mahdieh Soleymani Baghshah
arXiv AI
Sep 17

From Alignment to Synthesis: Contrastive Volumetric Grounding for Text-to-CT Generation

The paper introduces a 3D-CLIP encoder trained with structured hard negatives to improve vision‑language alignment for text‑to‑CT generation. This encoder drives a latent diffusion model that operates directly in 3D latent space, eliminating spatial artifacts from super‑resolution pipelines. Experiments on the CT‑RATE dataset show state‑of‑the‑art image fidelity and factual correctness across 18 pathological conditions, with lower inference time and GPU memory usage than competing methods.

By Daniele Molino, Camillo Maria Caruso, Filippo Ruffini, Paolo Soda, Valerio Guarrasi
arXiv AI
Jul 8

Few Channels Draw The Whole Picture: Revealing Massive Activations in Diffusion Transformers

arXiv:2605. 13974v2 Announce Type: replace-cross Abstract: Diffusion Transformers (DiTs) and related flow-based architectures are now among the strongest text-to-image generators, yet the internal mechanisms through which prompts shape image semantics remain poorly understood.

By Evelyn Turri, Davide Bucciarelli, Sara Sarto, Lorenzo Baraldi, Marcella Cornia
arXiv AI
Jun 24

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models

arXiv:2605. 08245v4 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) increasingly power high-stakes applications, from medical imaging to autonomous systems, yet they routinely hallucinate, confidently describing content not present in the input.

By Harshvardhan Saini, Samyak Jha, Yiming Tang, Dianbo Liu
arXiv AI
Oct 1

D-Scope: Decomposing and Steering Diffusion Transformers with Sparse Autoencoders

D‑Scope is a framework that links the interpretation of sparse autoencoder (SAE) features in diffusion transformers (DiTs) to controllable image generation. It aggregates SigLIP‑2 embeddings of highly activating image patches into visual centroids, matches target text descriptions against these centroids, and retrieves individual features without per‑feature text annotations. The method provides visual evidence for each selection and uses spatially masked interventions to test decoder directions under fixed generation conditions, evaluated across 150 SAEs and a benchmark of 100 target concepts.

By Xinyue Xu, Jiahao Zhang, Lijie Hu, Peter Hase, Hao Wang