arXiv Machine Learning By Arshia Hemmat, Amirhossein Vahidi, Amitis Shidani, Mohammad Vali Sanian, Hesam Asadollahzadeh, Aryan Yazdan Parast, Mohammad Lotfollahi

ORCA: Hunting Compositional Failures in Text-to-Image Diffusion

Read the original on arXiv Machine Learning →

ORCA (Orthogonal Residual Compositional Alignment) is a method that improves text-to-image diffusion models by aligning the diffusion transformer’s latent space with a low‑rank target derived from a frozen visual encoder. It introduces an auxiliary loss that uses a predictor with an orthogonal basis parameterised by a learned residual between T5 and CLIP embeddings, providing a prompt‑dependent signal for selecting the visual readout subspace. Experiments on three diffusion‑transformer backbones show that ORCA improves FID and GenEval scores, especially on attribute binding, spatial relations, and multi‑object prompts, without adding inference‑time cost.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computation and Language
Aug 28

Visual Information-Guided Parallel Decoding for Diffusion Multimodal Large Language Models

Visual Information-Guided Parallel Decoding for Diffusion Multimodal Large Language Models introduces the VIG‑Sampler, a method that prioritizes tokens for decoding based on their attention to image tokens and penalizes redundancy in image‑attention distributions. The approach aims to improve the quality of multimodal generation by selecting more informative tokens during diffusion decoding. Experiments on seven captioning and VQA benchmarks with three open‑source dMLLMs show that VIG‑Sampler outperforms the Info‑Gain Sampler by an average of 19.3 CIDEr points and achieves better COCO Caption results using only half as many decoding steps.

By Insu Lee, Wooje Park, Wonseok Shin, Jinwoo Son, Byonghyo Shim
arXiv Computer Vision
Aug 25

VISTA: Test-Time Compositional Alignment for Visual Autoregressive Generation

VISTA is a gradient‑based test‑time alignment framework designed for next‑scale visual autoregressive (VAR) image generation. It optimizes intermediate representations within the frozen transformer to enforce compositional constraints, without altering model weights or requiring extra training. Experiments on two benchmarks and two model scales show that VISTA improves compositional accuracy by up to 20% on a 2B backbone and 6% on an 8B backbone, while preserving image quality and enabling a smaller model to outperform a larger one.

By Hossein Shahabadi, Niki Sepasian, Mahdieh Soleymani Baghshah
arXiv AI
Sep 17

From Alignment to Synthesis: Contrastive Volumetric Grounding for Text-to-CT Generation

The paper introduces a 3D-CLIP encoder trained with structured hard negatives to improve vision‑language alignment for text‑to‑CT generation. This encoder drives a latent diffusion model that operates directly in 3D latent space, eliminating spatial artifacts from super‑resolution pipelines. Experiments on the CT‑RATE dataset show state‑of‑the‑art image fidelity and factual correctness across 18 pathological conditions, with lower inference time and GPU memory usage than competing methods.

By Daniele Molino, Camillo Maria Caruso, Filippo Ruffini, Paolo Soda, Valerio Guarrasi