arXiv:2607. 02235v1 Announce Type: cross Abstract: LLM-as-a-Judge has become the dominant evaluation paradigm for many natural language generation tasks, due to shortcomings of conventional metrics and high correlations with human judgment, albeit mostly in English.
By A. Seza Do\u{g}ru\"oz, Xixian Liao, Verena Blaschke, Jakob Prange, Senyu Li, David Ifeoluwa Adelani
arXiv:2607.02235v2 Announce Type: replace-cross
Abstract: LLM-as-a-Judge has become the dominant evaluation paradigm for many natural language generation tasks (albeit mostly in English) due to short...
By A. Seza Do\u{g}ru\"oz, Xixian Liao, Verena Blaschke, Jakob Prange, Senyu Li, David Ifeoluwa Adelani
The paper "What Do We Expect from LLMs? Mapping the Design of LLM Benchmarks" analyzes 14,767 arXiv submissions from 2022 to 2026 that introduce or update evaluation resources for large language models. It systematically maps changes in target systems, domains, evaluation materials, conditions, and scoring mechanisms, revealing a growing emphasis on action, interaction, and professional applications. The study also notes uneven development in model participation, with LLM-based scoring increasing in both agent and non-agent groups, while model-generated materials do not show a comparable rise.
By Chao Wang (Independent Researcher)
arXiv:2604. 09497v2 Announce Type: replace-cross Abstract: Accurate evaluation is central to the large language model (LLM) ecosystem, guiding model selection and downstream adoption across diverse use cases.
By Hippolyte Gisserot-Boukhlef, Nicolas Boizard, Emmanuel Malherbe, C\'eline Hudelot, Pierre Colombo
arXiv:2609.24516v1 Announce Type: new
Abstract: In recent years, large language models (LLMs) have emerged as a popular alternative for evaluation. Often referred to as LLMs as judges (LLJs), these s...
By Khaoula Chehbouni, Melina Medjdoub, Florian Carichon, Golnoosh Farnadi, Jackie Chi Kit Cheung
arXiv:2603. 23841v2 Announce Type: replace-cross Abstract: While Large Language Models (LLMs) are increasingly used as primary sources of information, their potential for political bias may impact their objectivity.
By Rohan Khetan, Ashna Khetan
arXiv:2607. 08731v2 Announce Type: replace-cross Abstract: National language models are becoming publicly funded epistemic infrastructure.
By Manuel Pita
arXiv:2607. 06940v1 Announce Type: cross Abstract: The remarkable performance of large language models (LLMs) in linguistic tasks underscores an urgent need for comprehensive evaluation of their response quality.
By Yiming Gai, Junde Lu, Xuefei Huang
The paper explores using large language models (LLMs) as AI respondents to convert policy documents into structured survey responses. It introduces a long-context in‑context learning pipeline that maps policy text to predefined survey categories such as policy instruments, target groups, and thematic areas, and includes a secondary LLM validation step. Evaluation on a multi‑country dataset shows high agreement (84‑95%) with human responses for structured indicators, though free‑text fields differ, indicating that hybrid human‑AI workflows can enhance policy monitoring efficiency while still requiring human oversight.
By Carolyn Cole, Matthias Deschryvere, Toqeer Ehsan, Arash Hajikhani
The study investigates whether large language models (LLMs) are more prone to errors when they doubt the plausibility of input data, a phenomenon termed context‑memory conflict. Using non‑English and low‑resource language datasets, the authors generate text from factual, counterfactual, and fictional RDF triples in English, Czech, Slovak, and Upper Sorbian, and evaluate faithfulness with both human annotations and an LLM judge (Kimi K3). Contrary to expectations, the results show only a weak context‑memory conflict: counterfactual inputs receive slightly lower faithfulness scores than factual ones, and the choice of LLM judge can significantly affect perceived conflict strength.
By Peter Kochelka, Ale\v{s} Manuel Pap\'a\v{c}ek, Vojt\v{e}ch Dvo\v{r}\'ak, Ond\v{r}ej Du\v{s}ek
arXiv:2604.19139v4 Announce Type: replace-cross
Abstract: Repeated praise, canned reassurance, familiar contrasts, and conspicuous vocabulary are recurring subjects in discussions of large language m...
By Shuai Wu, Xue Li, Zhijun Wang, Bolun Liu, Weilin Cai, Zihao Su, Ran Wang
A national language model offers a linguistic community its own instrument for measuring what its citizens say and value. Portugal's AMALIA, a publicly funded 9B-parameter model for European Portuguese, appears competitive on agreement alone: asked to code the moral foundation of authority, it agrees with trained human coders to within six F1 points of open models eight to thirteen times its size.