arXiv:2607. 16412v1 Announce Type: new Abstract: Current benchmarks for language models primarily evaluate execution on fully specified tasks.
By Andy Dai, Zexue He, Zhenyu Zhang, Alex Pentland, Jiaxin Pei
arXiv:2607. 21306v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as tutors and thought partners, helping users reason through problems.
By Verona Teo, Raghav Jain, Tobias Gerstenberg, Max Kleiman-Weiner
arXiv:2608. 12372v1 Announce Type: new Abstract: AI systems are increasingly employed as decision aids, decision delegates, or autonomous decision-makers.
By Vijay Keswani, Breanna K. Nguyen, Cyrus Cousins, Vincent Conitzer, Walter Sinnott-Armstrong, Jana Schaich Borg
arXiv:2605. 28882v2 Announce Type: replace-cross Abstract: With the rapid advancement of large language models, evaluating human-likeness in open-ended conversation has become increasingly important.
By Yihang Lin, Yunze Gao, Zeyang Lin, Dongbo Li, Kun Peng, Yue Liu
arXiv:2605. 17064v2 Announce Type: replace Abstract: Large language models are optimized for instruction following and agentic tasks remain poorly aligned with the requirements of high-quality creative writing.
By Jan Zierstek, Matteo Batelic, Maya Medjad, Tim Sch\"onenberger
arXiv:2608. 01366v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are integral to complex intellectual tasks, yet output quality remains constrained by user-provided prompts.
By B. Sankar, Pawni Yadav, Srinidhi Ranjini Girish, Amogh A. S