arXiv AI By Mingjia Li, Jin Wu, Hong Qian, Wenhao Huang, Yiyang Huang, Yiwen Zhang, Chanjin Zheng, Xiangfeng Wang, Aimin Zhou, Jiajun Guo

IntElicit: Eliciting and Assessing Contextualized Creativity via Dialogue Policy Optimization

Read the original on arXiv AI →

arXiv:2606. 12086v1 Announce Type: new Abstract: Contextualized assessment offers high ecological validity for evaluating creativity but introduces a critical challenge: observed performance may be confounded with cognitive proficiency (domain knowledge) and agency (willingness to engage).

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.

arXiv AI
Jun 30

The Human Creativity Benchmark

arXiv:2606. 30561v1 Announce Type: new Abstract: Modern AI evaluation frameworks treat evaluator disagreement as noise to be resolved.

By Aspen Hopkins, Allison Nulty, Alexandria Minetti, Anoop Pakki, Angad Singh