The paper "What Do We Expect from LLMs? Mapping the Design of LLM Benchmarks" analyzes 14,767 arXiv submissions from 2022 to 2026 that introduce or update evaluation resources for large language models. It systematically maps changes in target systems, domains, evaluation materials, conditions, and scoring mechanisms, revealing a growing emphasis on action, interaction, and professional applications. The study also notes uneven development in model participation, with LLM-based scoring increasing in both agent and non-agent groups, while model-generated materials do not show a comparable rise.
By Chao Wang (Independent Researcher)
The paper evaluates how well current Large Language Models can translate natural language goals, written by video game testers, into well‑formed PDDL targets for classical planning. Using a carefully designed prompt template, six state‑of‑the‑art LLMs were tested on correctness, speed, and error tendencies with real‑world benchmarks. All models achieved high correctness (>92%), with Gemini 2.5 Flash reaching 96% accuracy and the fewest false positives, while GPT‑4.1 was the fastest, yet differences in performance and occasional failures due to ambiguity and domain limits remain.
By Tomas Balyo, Lukas Chrpa, G. Michael Youngblood
The paper proposes using large language models (LLMs) to identify disagreements among models as a way to focus expert effort on revising codebooks for large‑scale text annotation. Three expert feedback methods are evaluated: editing LLM‑generated revisions (Codebook Verifying), answering questions about disagreements (Question Answering), and labeling disagreement cases with rationales (Rationale Labeling). Experiments on tutoring‑session transcripts show that Rationale Labeling achieves the highest LLM‑labeling accuracy (64.9%) compared to the expert‑revised codebook (57.8%), with Question Answering also outperforming the baseline (60.5%).
By Zeyu He, Zhuqian Zhou, Kirk Vanacore, Rene F. Kizilcec, Ting-Hao 'Kenneth' Huang
arXiv:2606. 10956v1 Announce Type: new Abstract: The deployment of Large Language Model (LLM) agents for computer automation is accelerating, yet their ability to navigate complex, professional-grade productivity software is largely untested.
By Tengchao Lv, Dongdong Zhang, Jiayu Ding, Yilin Jia, Yuzhong Zhao, Yupan Huang, Wenshan Wu, Xiangyang Zhou, Shaohan Huang, Nan Yang, Li Dong, Lei Cui, Furu Wei
arXiv:2609.24516v1 Announce Type: new
Abstract: In recent years, large language models (LLMs) have emerged as a popular alternative for evaluation. Often referred to as LLMs as judges (LLJs), these s...
By Khaoula Chehbouni, Melina Medjdoub, Florian Carichon, Golnoosh Farnadi, Jackie Chi Kit Cheung
arXiv:2608. 03803v1 Announce Type: cross Abstract: Multilingual language models are deployed across a hundred or more languages, yet most benchmarks test whether a model can perform a task _in_ a language rather than whether it commands the language itself, conflating fluency with proficiency.
By Tom\'a\v{s} Burkert, Angelika Peljak-{\L}api\'nska, David Zelen\'y