arXiv:2511.21415v2 Announce Type: replace
Abstract: We introduce DiverseVAR, a framework that enhances the diversity of text-conditioned visual autoregressive models (VAR) at test time without requir...
By Mingue Park, Prin Phunyaphibarn, Phillip Y. Lee, Minhyuk Sung
arXiv:2511.19811v2 Announce Type: replace-cross
Abstract: Image diversity remains a fundamental challenge for text-to-image diffusion models. Low-diversity generation often leads to repetitive output...
By Debin Meng, Chen Jin, Zheng Gao, Yanran Li, Ioannis Patras, Georgios Tzimiropoulos
arXiv:2603. 28762v2 Announce Type: replace-cross Abstract: Modern Text-to-Image (T2I) diffusion models have achieved remarkable semantic alignment, yet they often suffer from a significant lack of variety, converging on a narrow set of visual solutions for any given prompt.
By Omer Dahary, Benaya Koren, Daniel Garibi, Daniel Cohen-Or
arXiv:2511.20251v2 Announce Type: replace
Abstract: Modern text-to-image models produce impressive visual results from richly specified prompts, yet their behavior under long prompts remains insuffic...
By Bo-Kai Ruan, Teng-Fang Hsiao, Ling Lo, Yi-Lun Wu, Hong-Han Shuai
Subject-driven personalized text-to-image generation requires a pretrained diffusion model to acquire a specific subject from a few reference images while preserving subject identity, following novel text prompts, and maintaining sample diversity. Existing optimization-based methods instantiate subject adaptation through full fine-tuning, textual embedding optimization, or low-rank parameter updates; PaRa further constrains personalization from the perspective of parameter rank reduction.
arXiv:2606. 02453v1 Announce Type: cross Abstract: Despite the remarkable fidelity of generative models, they frequently suffer from mode collapse.
By Xiang Li, Dianbo Liu, Kenji Kawaguchi
Visual autoregressive (VAR) models have emerged as a fast, high-quality alternative to diffusion for text-to-image generation, but like diffusion models they exhibit persistent compositional failures,...
VISTA is a gradient‑based test‑time alignment framework designed for next‑scale visual autoregressive (VAR) image generation. It optimizes intermediate representations within the frozen transformer to enforce compositional constraints, without altering model weights or requiring extra training. Experiments on two benchmarks and two model scales show that VISTA improves compositional accuracy by up to 20% on a 2B backbone and 6% on an 8B backbone, while preserving image quality and enabling a smaller model to outperform a larger one.
By Hossein Shahabadi, Niki Sepasian, Mahdieh Soleymani Baghshah
arXiv:2511. 18050v1 Announce Type: cross Abstract: Diffusion transformers have recently delivered strong text-to-image generation around 1K resolution, but we show that extending them to native 4K across diverse aspect ratios exposes a tightly coupled failure mode spanning positional encoding, VAE compression, and optimization.
By Tian Ye, Song Fei, Lei Zhu
arXiv:2608.24293v1 Announce Type: new
Abstract: Latent diffusion models have emerged as a dominant framework for high-fidelity image and video synthesis, operating in compact latent spaces with varia...
By Yeonkyeong Lee, Hyunsung Go, Jongmin Kim, Sewoong Lim, Donghoon Lee
arXiv:2606. 17979v1 Announce Type: new Abstract: Existing RL post-training methods for text-to-image generation usually convert the final-image reward into a single scalar advantage and apply it with the same strength to the entire generative trajectory.
By Jinjie Shen, Wei Deng, Xian Hu, Daiguo Zhou, Jian Luan
arXiv:2512.23245v3 Announce Type: replace
Abstract: Recent text-to-image diffusion models have significantly improved visual quality and text alignment. However, generating a sequence of images while...
By Shin Seong Kim, Minjung Shin, Hyunin Cho, Youngjung Uh