The paper investigates whether large language models (LLMs) assess politeness in ways that match human judgments. Using two English datasets—one with continuous ratings and another with three‑way categorical labels—the authors find that LLMs agree more with each other than with humans. They observe that model–human alignment depends on explicit linguistic cues, while misaligned cases often involve rapport‑building strategies. Additionally, models tend to overproduce Neutral labels and underpredict Impolite labels, a pattern that persists even when expert consensus is used as a reference.
The paper introduces UPHELD, a large benchmark of human-to-human dialogues written by professional script writers, featuring realistic turn densities and over 36,000 per-turn human annotations. It evaluates existing automatic metrics and LLM-as-a-judge methods, finding them unreliable against expert human judgment. Using UPHELD, the authors develop a Mixture-of-Judges framework that improves correlation with human assessments by about 30%.
By Ilija Subasic, Andrew Rabinovich, Zhao Chen
arXiv:2608. 12368v1 Announce Type: new Abstract: Agreement with human judgments is a common proxy for evaluating the alignment of large language models (LLMs).
By Octavian M. Machidon, Alina L. Machidon, Vojko Strahovnik, Mateja Centa Strahovnik, Jonas Miklav\v{c}i\v{c}, Marko Robnik \v{S}ikonja
arXiv:2606. 12754v1 Announce Type: cross Abstract: Are large language models (LLMs) bad at capturing human judgment?
By Danica Dillion, Chen Cecilia Liu, Baihui Wang, Daniele Barolo, Tanmay Rajore, Niket Tandon, Pranathi Ravikumar, Kurt Gray
arXiv:2604. 02512v2 Announce Type: replace-cross Abstract: Large language models (LLMs) increasingly exhibit human-like patterns of pragmatic and social reasoning.
By Roland M\"uhlenbernd
PADM'E is a method for synthesizing preference‑aligned data to meta‑evaluate language‑model (LM) evaluators of agentic behaviors. It reframes meta‑evaluation as a preference judgment problem, generating criterion‑based data with small LMs and no human involvement. In a prototype, PADM'E produced 1,000 samples across four domains and three criteria, and human validation showed agreement with human judgment rising from 73% to 85% compared to a naive baseline.
By Cheng Chang, Yining Mao, Peng Qi