arXiv AI By Aspen Hopkins, Allison Nulty, Alexandria Minetti, Anoop Pakki, Angad Singh

The Human Creativity Benchmark

Read the original on arXiv AI →

arXiv:2606. 30561v1 Announce Type: new Abstract: Modern AI evaluation frameworks treat evaluator disagreement as noise to be resolved.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 26

The Limits of Automatic Evaluation of Creativity in Large Language Models

The paper examines whether existing automatic methods can reliably assess creativity in text produced by large language models (LLMs). By collecting human ratings on 11 creativity dimensions for both human and AI short stories, the authors compare these judgments with automated metrics and LLM-as-a-Judge evaluations. The results show a significant misalignment: automated metrics and LLM judges favor AI-generated stories and show near-zero correlation with human assessments, revealing fundamental limitations in current computational approaches to evaluating creative text.

By Alessandro Tutone, Giorgio Franceschelli, Mirco Musolesi
arXiv Computation and Language
Sep 4

Decoupled Analysis-Judging: An Automated Creativity Evaluator Using LLMs in Complex Multi-step Creativity Tasks

The paper introduces CreaEval, an automated creativity evaluator designed for complex multi-step tasks (CGPST). It separates the evaluation process into two phases: a memory‑augmented analysis that transforms responses into structured evidence, and an evidence‑based judging step that scores without seeing raw outputs. Experiments show CreaEval outperforms existing baselines by an average of 22.74% across CGPST and two simpler creativity tasks.

By Xiangyu Wang, Jin Wu, Xiaoyu Li, Chanjin Zheng, Yifeng Zhou
arXiv AI
Jun 11

IntElicit: Eliciting and Assessing Contextualized Creativity via Dialogue Policy Optimization

arXiv:2606. 12086v1 Announce Type: new Abstract: Contextualized assessment offers high ecological validity for evaluating creativity but introduces a critical challenge: observed performance may be confounded with cognitive proficiency (domain knowledge) and agency (willingness to engage).

By Mingjia Li, Jin Wu, Hong Qian, Wenhao Huang, Yiyang Huang, Yiwen Zhang, Chanjin Zheng, Xiangfeng Wang, Aimin Zhou, Jiajun Guo