arXiv AI By Mohammad Jalali, Azim Ospanov, Amin Gohari, Farzan Farnia

Conditional Vendi Score: Prompt-Aware Diversity Evaluation for Generative AI Models and LLMs

Read the original on arXiv AI →

arXiv:2411. 02817v2 Announce Type: replace-cross Abstract: Generative models guided by text prompts are widely evaluated for fidelity and prompt alignment, yet their ability to produce outputs remains underexplored.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 3

Do Large Language Models Capture the Diversity in their Training Data?

The paper investigates whether large language models (LLMs) capture the full diversity of outputs present in their training data. Using an information‑theoretic approach, the authors compare the conditional entropy of model‑generated outputs with that of the training data, finding that LLMs consistently produce outputs with lower conditional entropy across various models, scales, and decoding strategies. They also extend the analysis to image and text‑conditioned generators, propose a post‑hoc correction method based on matrix‑entropy projection to increase conditional diversity, and provide theoretical guarantees and an efficient algorithm for this correction.

By Youqi Wu, Farzan Farnia
arXiv Computer Vision
Sep 3

Diversifying Long Prompt Image Generation through Structured Prompt Embedding Space Sampling

The paper investigates how long, richly detailed prompts cause modern text-to-image models to lose diversity, even when many visual aspects are unspecified. It introduces PromptMoG, a training‑free method that samples prompt embeddings from a Mixture‑of‑Gaussians distribution to restore diversity while preserving semantic fidelity. The authors also present LPD‑Bench, a benchmark for evaluating fidelity and diversity under long, semantically dense prompts, and demonstrate PromptMoG’s effectiveness on four large diffusion models.

By Bo-Kai Ruan, Teng-Fang Hsiao, Ling Lo, Yi-Lun Wu, Hong-Han Shuai
arXiv AI
Aug 19

Evaluating the Diversity of AI-Generated Content with Diversity Profiles

The paper argues that measuring diversity in AI-generated content using a single scalar score is inherently ambiguous and often misleading. It reviews existing diversity metrics, demonstrates their limitations through axiomatic and empirical analyses, and introduces diversity profiles—curve-valued, condition-aware summaries that evaluate diversity across a range of thresholds, scales, exponents, or orders. These profiles reveal whether comparisons are robust across resolutions or depend on arbitrary parameter choices, offering a more transparent framework for generative AI evaluation.

By Xiuyuan Hu, Xuege Hou, Guoqing Liu, Yang Zhao, Jieran Li, Dongbiao Sun, Jos\'e Miguel Hern\'andez-Lobato, Hao Zhang, Xue Liu