CreativityPrism: A Cross-Domain Evaluation Framework for Large Language Model Creativity
arXiv:2510. 20091v3 Announce Type: replace-cross Abstract: Creativity is often seen as a hallmark of human intelligence.
arXiv:2604. 03374v2 Announce Type: replace-cross Abstract: Creative problem-solving requires combining multiple cognitive abilities, including logical reasoning, lateral thinking, analogy-making, and commonsense knowledge, to discover insights that connect seemingly unrelated pieces of information.
arXiv:2510. 20091v3 Announce Type: replace-cross Abstract: Creativity is often seen as a hallmark of human intelligence.
arXiv:2606. 11762v1 Announce Type: cross Abstract: Large language models (LLMs) have achieved remarkable progress in language understanding, reasoning, and generation, sparking growing interest in their creative potential.
arXiv:2605. 10574v3 Announce Type: replace Abstract: As artificial intelligence advances, models are not improving uniformly.
Large language models (LLMs) have achieved remarkable progress in language understanding, reasoning, and generation, sparking growing interest in their creative potential. Realizing this potential requires systematic and scalable methods for evaluating creativity across diverse tasks.
arXiv:2608. 06501v1 Announce Type: new Abstract: Creative capabilities of MLLMs matter in design, communication, education, and human--AI collaboration, yet remain difficult to evaluate because explicit targets and reward signals are scarce compared with accuracy-oriented tasks.
arXiv:2607. 14109v1 Announce Type: cross Abstract: Probing the capabilities of Large Language Models (LLMs) and building robust solutions for Multiple-Choice Question Answering (MCQA) remain central challenges in natural language understanding.
arXiv:2509. 14704v3 Announce Type: replace Abstract: Benchmark saturation and training-data contamination increasingly obscure whether reported gains in large language models (LLMs) reflect genuine advances in reasoning or familiarity with recurring patterns in benchmark problems.
WordPolo is a word‑finding task that evaluates language models by having them guess an unknown target word and receive semantic similarity feedback. Participants start with no knowledge, make iterative guesses, and receive distance scores that guide them through semantic space. The study tests recent LLMs, LRMs, humans, and a heuristic on 1,500 puzzles, revealing that while solve rates vary widely, many models make meaningful progress and exhibit human‑like strategies, highlighting the importance of assessing reasoning processes, not just final accuracy.
arXiv:2605. 03344v2 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) has proven effective for knowledge-intensive tasks, but is widely believed to offer limited benefit for reasoning-intensive problems such as math and code generation.
arXiv:2609.38406v1 Announce Type: new Abstract: Access to real-world information is often noisy and fragmented. Constructing a coherent narrative from such fragments requires models to reconstruct mi...
The paper introduces CreaEval, an automated creativity evaluator designed for complex multi-step tasks (CGPST). It separates the evaluation process into two phases: a memory‑augmented analysis that transforms responses into structured evidence, and an evidence‑based judging step that scores without seeing raw outputs. Experiments show CreaEval outperforms existing baselines by an average of 22.74% across CGPST and two simpler creativity tasks.
arXiv:2510. 12171v2 Announce Type: replace Abstract: Large Language Models have shown strong scientific reasoning ability, but their performance on materials science problems remains less studied.