arXiv:2606. 30561v1 Announce Type: new Abstract: Modern AI evaluation frameworks treat evaluator disagreement as noise to be resolved.
By Aspen Hopkins, Allison Nulty, Alexandria Minetti, Anoop Pakki, Angad Singh
arXiv:2510. 20091v3 Announce Type: replace-cross Abstract: Creativity is often seen as a hallmark of human intelligence.
By Zhaoyi Joey Hou, Bowei Alvin Zhang, Yining Lu, Bhiman Kumar Baghel, Anneliese Brei, Ximing Lu, Meng Jiang, Faeze Brahman, Snigdha Chaturvedi, Haw-Shiuan Chang, Daniel Khashabi, Xiang Lorraine Li
Creativity is a complex cognitive ability that relies on knowledge organisation and retrieval from semantic memory. Yet most research uses a single task to measure it, capturing only a fraction of this complexity.
AesCanvas is a new dataset and benchmark that evaluates image aesthetic models on two fronts: CritiqueCanvas, which contains 519,136 instruction–response pairs for long‑form, multi‑dimensional critique across photography, painting, and virtual imagery, and ContextCanvas, which offers 301 expert‑reviewed use scenarios to assess contextual aesthetic suitability. The benchmark tests closed‑source, open‑weight general, and aesthetic‑specific multimodal large language models, revealing that models excel at critique generation but lag in context‑sensitive judgment. The study shows that aesthetic specialization does not reliably transfer to contextual suitability and highlights the need for culturally situated, evidence‑grounded suitability as a distinct objective for aesthetic modeling.
By Xuanwei Hu, Haoyu Dong, Kejun Wu, Tianyi Liu, Jianjun Gao
arXiv:2603. 19087v2 Announce Type: replace Abstract: Creativity is the ability to come up with novel ideas, a capacity crucial for human development and flourishing.
By Qiawen Ella Liu, Marina Dubova, Henry Conklin, Takumi Harada, Thomas L. Griffiths
The paper examines whether existing automatic methods can reliably assess creativity in text produced by large language models (LLMs). By collecting human ratings on 11 creativity dimensions for both human and AI short stories, the authors compare these judgments with automated metrics and LLM-as-a-Judge evaluations. The results show a significant misalignment: automated metrics and LLM judges favor AI-generated stories and show near-zero correlation with human assessments, revealing fundamental limitations in current computational approaches to evaluating creative text.
By Alessandro Tutone, Giorgio Franceschelli, Mirco Musolesi