The paper investigates whether large language models (LLMs) capture the full diversity of outputs present in their training data. Using an information‑theoretic approach, the authors compare the conditional entropy of model‑generated outputs with that of the training data, finding that LLMs consistently produce outputs with lower conditional entropy across various models, scales, and decoding strategies. They also extend the analysis to image and text‑conditioned generators, propose a post‑hoc correction method based on matrix‑entropy projection to increase conditional diversity, and provide theoretical guarantees and an efficient algorithm for this correction.
By Youqi Wu, Farzan Farnia
The paper investigates how long, richly detailed prompts cause modern text-to-image models to lose diversity, even when many visual aspects are unspecified. It introduces PromptMoG, a training‑free method that samples prompt embeddings from a Mixture‑of‑Gaussians distribution to restore diversity while preserving semantic fidelity. The authors also present LPD‑Bench, a benchmark for evaluating fidelity and diversity under long, semantically dense prompts, and demonstrate PromptMoG’s effectiveness on four large diffusion models.
By Bo-Kai Ruan, Teng-Fang Hsiao, Ling Lo, Yi-Lun Wu, Hong-Han Shuai
The paper argues that measuring diversity in AI-generated content using a single scalar score is inherently ambiguous and often misleading. It reviews existing diversity metrics, demonstrates their limitations through axiomatic and empirical analyses, and introduces diversity profiles—curve-valued, condition-aware summaries that evaluate diversity across a range of thresholds, scales, exponents, or orders. These profiles reveal whether comparisons are robust across resolutions or depend on arbitrary parameter choices, offering a more transparent framework for generative AI evaluation.
By Xiuyuan Hu, Xuege Hou, Guoqing Liu, Yang Zhao, Jieran Li, Dongbiao Sun, Jos\'e Miguel Hern\'andez-Lobato, Hao Zhang, Xue Liu
arXiv:2608. 09385v1 Announce Type: cross Abstract: Generative AI models are primarily designed to imitate the data distribution, an objective that neither corrects diversity lost by a learned generator nor defines how generation should extend beyond the diversity of the data itself.
By Hossein Goli, Farzan Farnia, Amin Gohari
arXiv:2307.00852v3 Announce Type: replace
Abstract: The natural language generation domain has witnessed great success thanks to Transformer models. Although they have achieved state-of-the-art gener...
By Yueen Ma, Dafeng Chi, Jingjing Li, Kai Song, Yuzheng Zhuang, Irwin King
arXiv:2608.29335v1 Announce Type: new
Abstract: Latent generative models typically follow a two-stage pipeline, training a variational autoencoder for reconstruction and then a generative model on th...
By Guangting Zheng, Yiyuan Zhang, Tao Yang, Yunpeng Chen, Rui Zhu, Jiajun Deng, Yanyong Zhang