arXiv Machine Learning

MaskAttn-SDXL: Controllable Region-Level Text-To-Image Generation

arXiv:2509. 15357v3 Announce Type: replace-cross Abstract: Diffusion models have achieved strong results in text-to-image generation, but important limitations remain as prompts become more structured and multi-object.

arXiv Computation and Language
Sep 14

Representation-based Masked Diffusion Model

The paper introduces Representation-based Masked Diffusion Model (RMDM), a new framework for language modeling that improves upon existing Masked Diffusion Models by incorporating global semantic guidance. RMDM encodes text into a continuous semantic space with a pretrained encoder, normalizes this representation to a Gaussian prior via an invertible transformation, and then trains a masked diffusion model conditioned on this latent representation to coordinate parallel token updates. Experiments show that RMDM yields higher generation quality, especially when using aggressive few‑step sampling.

By Yangrong Hu, Ding Huang, Xueyu Zhou, Jian Huang
arXiv Computer Vision
Aug 25

VISTA: Test-Time Compositional Alignment for Visual Autoregressive Generation

VISTA is a gradient‑based test‑time alignment framework designed for next‑scale visual autoregressive (VAR) image generation. It optimizes intermediate representations within the frozen transformer to enforce compositional constraints, without altering model weights or requiring extra training. Experiments on two benchmarks and two model scales show that VISTA improves compositional accuracy by up to 20% on a 2B backbone and 6% on an 8B backbone, while preserving image quality and enabling a smaller model to outperform a larger one.

By Hossein Shahabadi, Niki Sepasian, Mahdieh Soleymani Baghshah
arXiv Computation and Language
Aug 28

Visual Information-Guided Parallel Decoding for Diffusion Multimodal Large Language Models

Visual Information-Guided Parallel Decoding for Diffusion Multimodal Large Language Models introduces the VIG‑Sampler, a method that prioritizes tokens for decoding based on their attention to image tokens and penalizes redundancy in image‑attention distributions. The approach aims to improve the quality of multimodal generation by selecting more informative tokens during diffusion decoding. Experiments on seven captioning and VQA benchmarks with three open‑source dMLLMs show that VIG‑Sampler outperforms the Info‑Gain Sampler by an average of 19.3 CIDEr points and achieves better COCO Caption results using only half as many decoding steps.

By Insu Lee, Wooje Park, Wonseok Shin, Jinwoo Son, Byonghyo Shim