arXiv AI

Recipes for Creativity: Iterative Generation and Evaluation in Large Language Models

arXiv:2608. 07243v1 Announce Type: new Abstract: Generative models are often evaluated through singular artifacts, whereas human creativity typically emerges through iterative generation, appraisal, and refinement.

arXiv AI
Aug 26

The Limits of Automatic Evaluation of Creativity in Large Language Models

The paper examines whether existing automatic methods can reliably assess creativity in text produced by large language models (LLMs). By collecting human ratings on 11 creativity dimensions for both human and AI short stories, the authors compare these judgments with automated metrics and LLM-as-a-Judge evaluations. The results show a significant misalignment: automated metrics and LLM judges favor AI-generated stories and show near-zero correlation with human assessments, revealing fundamental limitations in current computational approaches to evaluating creative text.

By Alessandro Tutone, Giorgio Franceschelli, Mirco Musolesi
arXiv AI
Jun 12

CreativeBench: Benchmarking and Enhancing Machine Creativity via Self-Evolving Challenges

arXiv:2603. 11863v2 Announce Type: replace Abstract: The saturation of high-quality pre-training data has shifted research focus toward evolutionary systems capable of continuously generating novel artifacts, leading to the success of AlphaEvolve.

By Zi-Han Wang, Lam Nguyen, Zhengyang Zhao, Mengyue Yang, Chengwei Qin, Yujiu Yang, Linyi Yang
arXiv AI
Jun 30

The Human Creativity Benchmark

arXiv:2606. 30561v1 Announce Type: new Abstract: Modern AI evaluation frameworks treat evaluator disagreement as noise to be resolved.

By Aspen Hopkins, Allison Nulty, Alexandria Minetti, Anoop Pakki, Angad Singh
arXiv AI
6d ago

Breaking Homogeneity: Diversifying Persona Sets for Creative LLM Outputs

The paper introduces persona diversification as a set‑level conditioning strategy to reduce homogeneity in large language model outputs. It explores two design axes—selecting versus generating personas and space‑filling versus frontier‑seeking diversity—implementing four methods that span coverage and dispersion subset selections, uniform‑coverage sampling, and evolutionary persona generation. Experiments on tasks such as the Alternative Uses Task, Infinity‑Chat, and Divergent Association Task demonstrate significant gains in response diversity, originality, flexibility, and overall creativity, while maintaining high validity.

By Sang Bin Moon, Nicole Cho, Daniel Borrajo, Sumitra Ganesh, Abolfazl Hashemi