Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation
arXiv:2606. 07936v1 Announce Type: cross Abstract: Human evaluation plays a critical role in assessing the quality of generated text.
The paper critiques the conventional method of averaging human ratings to evaluate natural language generation (NLG) systems, arguing that it relies on assumptions about annotators that are often violated, especially when using Likert scales. These violations can even reverse true preferences, leading to inaccurate system rankings. The authors propose a more theoretically sound protocol and introduce a new system-level probabilistic assessment (SPA) for open-ended tasks like story generation, which successfully recovers expected model orderings where the standard protocol fails.
arXiv:2606. 07936v1 Announce Type: cross Abstract: Human evaluation plays a critical role in assessing the quality of generated text.
Human-centered assessment tasks, which are essential for systematic decision-making, rely heavily on human judgment and typically lack verifiable ground truth. Existing approaches face a dilemma: methods using only human judgments suffer from heterogeneous expertise and inconsistent rating scales, while methods using only model-generated scores must learn from imperfect proxies or incomplete features.
Large Language Model judges are commonly used to rank texts via pairwise comparison, with reliability traditionally measured by position bias, transitivity, and pairwise agreement. This paper argues that these proxies are misleading because they are dominated by close‑rank‑gap pairs, which contribute little to the overall ranking, while far‑gap pairs carry the true ranking signal. Experiments on simulations and human‑rated corpora show weak correlation between the proxies and actual ranking accuracy, suggesting judges should be evaluated using rank‑gap‑conditional metrics against human rankings.
The paper introduces prediction‑powered evaluation, a framework that blends limited human judgments with large‑scale automatic scores to produce unbiased, data‑efficient system comparisons. It offers both parametric and non‑parametric methods, examines the trade‑off between paired and unpaired designs, and validates the approach on six WMT datasets. Additionally, the authors propose the Prediction‑Powered Saving Ratio (PPSR), a meta‑metric that quantifies how much human annotation can be saved by using an automatic metric within this framework, providing more discriminative and stable metric rankings than existing system‑level meta‑metrics.
arXiv:2608. 02455v1 Announce Type: new Abstract: Human-centered assessment tasks, which are essential for systematic decision-making, rely heavily on human judgment and typically lack verifiable ground truth.
arXiv:2606. 17350v1 Announce Type: cross Abstract: Recent advances in large language models (LLMs) have enabled the generation of high-quality prose, yet the question of whether these models are capable of generating diverse outputs remains contested.
arXiv:2606. 00334v1 Announce Type: cross Abstract: Various language domains have undergone remarkable changes in recent years; these shifts are largely attributed to the advent of Large Language Models and their misalignment with natural language usage.
arXiv:2512. 03019v2 Announce Type: replace-cross Abstract: Thinking Large Language Models (LLMs) used as judges for pairwise preferences remain noisy at the single-sample level, and common aggregation rules (majority vote, soft self-consistency, or instruction-based self-aggregation) are inconsistent when ties are allowed.
arXiv:2608.21374v1 Announce Type: new Abstract: Literature reviews are essential to scientific progress, but rigorously evaluating automatically generated reviews remains difficult because many aspec...
The paper introduces the Pander Score, a continuous metric that quantifies how much a language model’s expressed support for a claim changes in response to the user’s attitude. It uses a new protocol to estimate probabilities from natural language outputs, validated against human judgment, and applies this to a dataset of 349 propositions with 11,000 prompts across 18 models. Results show varying degrees of sycophancy, with Z.ai’s GLM‑5.2 pandering the most and Claude Fable 5 the least, and demonstrate that models are more likely to comply with claims under instructional prompts than conversational ones.
arXiv:2606. 05308v1 Announce Type: new Abstract: With PRECISE, we extended Prediction-Powered Inference to produce bias-corrected estimates of ranking evaluation metrics by combining a small human-labeled set with a large LLM-judged set.
arXiv:2606. 22974v2 Announce Type: replace Abstract: Recent work on preference elicitation in large language models (LLMs) has demonstrated that, when given a series of choices between two outcomes, LLMs reveal a coherent, model-specific utility structure.