The paper examines whether existing automatic methods can reliably assess creativity in text produced by large language models (LLMs). By collecting human ratings on 11 creativity dimensions for both human and AI short stories, the authors compare these judgments with automated metrics and LLM-as-a-Judge evaluations. The results show a significant misalignment: automated metrics and LLM judges favor AI-generated stories and show near-zero correlation with human assessments, revealing fundamental limitations in current computational approaches to evaluating creative text.
By Alessandro Tutone, Giorgio Franceschelli, Mirco Musolesi
arXiv:2510. 20091v3 Announce Type: replace-cross Abstract: Creativity is often seen as a hallmark of human intelligence.
By Zhaoyi Joey Hou, Bowei Alvin Zhang, Yining Lu, Bhiman Kumar Baghel, Anneliese Brei, Ximing Lu, Meng Jiang, Faeze Brahman, Snigdha Chaturvedi, Haw-Shiuan Chang, Daniel Khashabi, Xiang Lorraine Li
arXiv:2607. 23696v1 Announce Type: new Abstract: Ad creative optimization is increasingly constrained by evaluation rather than generation.
By Kevin Lee, Benjamin Letham, Zhiyuan Jerry Lin, Elodie Samson, Eric Onofrey, Poppy Zhang, Shawndra Hill, Eytan Bakshy
arXiv:2608. 19437v1 Announce Type: cross Abstract: Many benchmarks track Large Language Model (LLM) performance on tasks with verifiable answers, but less is known about how LLM performance is evolving on open-ended tasks, where creativity, originality and diversity may matter as much as quality.
By Nirav Patel, Josiah Crossman, Eva Aggarwal, Emily Wenger
arXiv:2603. 11863v2 Announce Type: replace Abstract: The saturation of high-quality pre-training data has shifted research focus toward evolutionary systems capable of continuously generating novel artifacts, leading to the success of AlphaEvolve.
By Zi-Han Wang, Lam Nguyen, Zhengyang Zhao, Mengyue Yang, Chengwei Qin, Yujiu Yang, Linyi Yang
Large language models (LLMs) have achieved remarkable progress in language understanding, reasoning, and generation, sparking growing interest in their creative potential. Realizing this potential requires systematic and scalable methods for evaluating creativity across diverse tasks.
arXiv:2606. 11762v1 Announce Type: cross Abstract: Large language models (LLMs) have achieved remarkable progress in language understanding, reasoning, and generation, sparking growing interest in their creative potential.
By Min Sen Tan, Zachary Kit Chun Choy, Syed Ali Redha Alsagoff, Nadya Yuki Wangsajaya, Mohor Banerjee, Swaagat Bikash Saikia, Alvin Chan
arXiv:2502. 13207v4 Announce Type: replace-cross Abstract: Despite the increasing use of large language models for creative tasks, their outputs often lack diversity.
By Giorgio Franceschelli, Mirco Musolesi
arXiv:2608.30047v1 Announce Type: new
Abstract: Recent AI systems promise autonomous scientific discovery, claiming to discover algorithms and produce research papers, yet understanding whether they...
By Shitanshu Bhushan, Yunxiang Zhang, Lu Wang
arXiv:2606. 30561v1 Announce Type: new Abstract: Modern AI evaluation frameworks treat evaluator disagreement as noise to be resolved.
By Aspen Hopkins, Allison Nulty, Alexandria Minetti, Anoop Pakki, Angad Singh
arXiv:2603. 19087v2 Announce Type: replace Abstract: Creativity is the ability to come up with novel ideas, a capacity crucial for human development and flourishing.
By Qiawen Ella Liu, Marina Dubova, Henry Conklin, Takumi Harada, Thomas L. Griffiths
The paper introduces persona diversification as a set‑level conditioning strategy to reduce homogeneity in large language model outputs. It explores two design axes—selecting versus generating personas and space‑filling versus frontier‑seeking diversity—implementing four methods that span coverage and dispersion subset selections, uniform‑coverage sampling, and evolutionary persona generation. Experiments on tasks such as the Alternative Uses Task, Infinity‑Chat, and Divergent Association Task demonstrate significant gains in response diversity, originality, flexibility, and overall creativity, while maintaining high validity.
By Sang Bin Moon, Nicole Cho, Daniel Borrajo, Sumitra Ganesh, Abolfazl Hashemi