The Human Creativity Benchmark
arXiv:2606. 30561v1 Announce Type: new Abstract: Modern AI evaluation frameworks treat evaluator disagreement as noise to be resolved.
arXiv:2607. 28644v1 Announce Type: cross Abstract: Creativity in computational systems is often evaluated as an objective property of artifacts, with existing Computational Creativity (CC) frameworks assessing creative merit at the level of outputs or systems rather than interpretive context.
arXiv:2606. 30561v1 Announce Type: new Abstract: Modern AI evaluation frameworks treat evaluator disagreement as noise to be resolved.
arXiv:2510. 20091v3 Announce Type: replace-cross Abstract: Creativity is often seen as a hallmark of human intelligence.
Creativity is a complex cognitive ability that relies on knowledge organisation and retrieval from semantic memory. Yet most research uses a single task to measure it, capturing only a fraction of this complexity.
AesCanvas is a new dataset and benchmark that evaluates image aesthetic models on two fronts: CritiqueCanvas, which contains 519,136 instruction–response pairs for long‑form, multi‑dimensional critique across photography, painting, and virtual imagery, and ContextCanvas, which offers 301 expert‑reviewed use scenarios to assess contextual aesthetic suitability. The benchmark tests closed‑source, open‑weight general, and aesthetic‑specific multimodal large language models, revealing that models excel at critique generation but lag in context‑sensitive judgment. The study shows that aesthetic specialization does not reliably transfer to contextual suitability and highlights the need for culturally situated, evidence‑grounded suitability as a distinct objective for aesthetic modeling.
arXiv:2603. 19087v2 Announce Type: replace Abstract: Creativity is the ability to come up with novel ideas, a capacity crucial for human development and flourishing.
The paper examines whether existing automatic methods can reliably assess creativity in text produced by large language models (LLMs). By collecting human ratings on 11 creativity dimensions for both human and AI short stories, the authors compare these judgments with automated metrics and LLM-as-a-Judge evaluations. The results show a significant misalignment: automated metrics and LLM judges favor AI-generated stories and show near-zero correlation with human assessments, revealing fundamental limitations in current computational approaches to evaluating creative text.
arXiv:2606. 11762v1 Announce Type: cross Abstract: Large language models (LLMs) have achieved remarkable progress in language understanding, reasoning, and generation, sparking growing interest in their creative potential.
arXiv:2608.10706v3 Announce Type: replace Abstract: Recent vision-language models demonstrate impressive general visual understanding, yet their art interpretation remains shallow: they describe surf...
arXiv:2608. 07500v1 Announce Type: cross Abstract: Research on human-GenAI collaboration yields conflicting findings: GenAI can enhance creativity yet reduce collective diversity, with uneven benefits across skill levels.
arXiv:2608. 14405v1 Announce Type: cross Abstract: Art style is a signature of professional digital artists that develops through repeated experimentation, reflection, and adaptation.
arXiv:2608. 19437v1 Announce Type: cross Abstract: Many benchmarks track Large Language Model (LLM) performance on tasks with verifiable answers, but less is known about how LLM performance is evolving on open-ended tasks, where creativity, originality and diversity may matter as much as quality.
arXiv:2608.30754v1 Announce Type: new Abstract: Evaluating creativity in large language model (LLM) outputs remains challenging because creativity is multidimensional and human-centered. We examine h...