arXiv:2608. 07243v1 Announce Type: new Abstract: Generative models are often evaluated through singular artifacts, whereas human creativity typically emerges through iterative generation, appraisal, and refinement.
By Rens Anderson, Tessa Verhoef, Amirhossein Zohrehvand
arXiv:2603. 11863v2 Announce Type: replace Abstract: The saturation of high-quality pre-training data has shifted research focus toward evolutionary systems capable of continuously generating novel artifacts, leading to the success of AlphaEvolve.
By Zi-Han Wang, Lam Nguyen, Zhengyang Zhao, Mengyue Yang, Chengwei Qin, Yujiu Yang, Linyi Yang
The paper introduces CreaEval, an automated creativity evaluator designed for complex multi-step tasks (CGPST). It separates the evaluation process into two phases: a memory‑augmented analysis that transforms responses into structured evidence, and an evidence‑based judging step that scores without seeing raw outputs. Experiments show CreaEval outperforms existing baselines by an average of 22.74% across CGPST and two simpler creativity tasks.
By Xiangyu Wang, Jin Wu, Xiaoyu Li, Chanjin Zheng, Yifeng Zhou
arXiv:2606. 11762v1 Announce Type: cross Abstract: Large language models (LLMs) have achieved remarkable progress in language understanding, reasoning, and generation, sparking growing interest in their creative potential.
By Min Sen Tan, Zachary Kit Chun Choy, Syed Ali Redha Alsagoff, Nadya Yuki Wangsajaya, Mohor Banerjee, Swaagat Bikash Saikia, Alvin Chan
Large language models (LLMs) have achieved remarkable progress in language understanding, reasoning, and generation, sparking growing interest in their creative potential. Realizing this potential requires systematic and scalable methods for evaluating creativity across diverse tasks.
arXiv:2510. 20091v3 Announce Type: replace-cross Abstract: Creativity is often seen as a hallmark of human intelligence.
By Zhaoyi Joey Hou, Bowei Alvin Zhang, Yining Lu, Bhiman Kumar Baghel, Anneliese Brei, Ximing Lu, Meng Jiang, Faeze Brahman, Snigdha Chaturvedi, Haw-Shiuan Chang, Daniel Khashabi, Xiang Lorraine Li
arXiv:2608.30047v1 Announce Type: new
Abstract: Recent AI systems promise autonomous scientific discovery, claiming to discover algorithms and produce research papers, yet understanding whether they...
By Shitanshu Bhushan, Yunxiang Zhang, Lu Wang
arXiv:2608. 07460v1 Announce Type: cross Abstract: While post-training improves the capabilities of large language models (LLMs), it generally lowers their output diversity and creativity, negatively impacting tasks that explicitly require creativity (e.
By Ananya Sahu, Mohit Bansal, Elias Stengel-Eskin
PinDCO is a scalable dynamic creative optimization system designed for Pinterest’s billion‑scale visual discovery platform. It uses a Creative Component Fusion Network to score ad creatives by modeling individual components (image, title, layout) with dedicated towers and fusing their representations, while a Pixel‑aware Adjustment Module tailors scores to creative size for better whole‑page outcomes. The system incorporates a lightweight pre‑selection model, caching, and dynamic batching to handle large candidate volumes, achieving a 3.09% lift in ad click‑through rate in online experiments.
By Yu Hao, Yuchun Li, Peimeng Sui, Meilin Liu, Tianyuan Cui, Hao Li, Zicong Zhou, Akanksha Baid
arXiv:2605. 28556v2 Announce Type: replace Abstract: As agent capabilities advance, existing benchmarks, such as $\tau^2$-Bench, are becoming increasingly saturated.
By Tomer Keren, Nitay Calderon, Asaf Yehudai, Yotam Perlitz, Michal Shmueli-Scheuer, Roi Reichart
arXiv:2606. 02863v1 Announce Type: new Abstract: AI-Driven Research Systems (ADRS) -- systems coupling LLMs with automated evaluation to discover algorithms, proofs, and designs -- are being optimized and adopted across domains, but the tools to analyze them have not kept pace.
By Marquita Ellis, Paul Castro
arXiv:2607. 29241v1 Announce Type: cross Abstract: Optimizing modern recommender models still depends heavily on engineers manually iterating over architectural, objective, and training-strategy changes.
By Haoran Ling, Yuecheng Li, Zeyu Song, Jing Yao, Shuwen Kang, Chi Lu, Wenjin Wu, Peng Jiang