When Consistency Does Not Mean Reliability: Evaluating Local LLM Judges Against Human Ratings
Read the original on arXiv Computation and Language →The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The Flow has not summarised this story yet — read it at arXiv Computation and Language.
arXiv:2601.08654v3 Announce Type: replace Abstract: Rubric-based text evaluation increasingly relies on large language models (LLMs) as scalable judges, yet frozen black-box models can interpret the...
arXiv:2609.13773v1 Announce Type: new Abstract: LLM-as-a-Judge evaluators are increasingly used to score open-ended generation, yet a judge's correlation with human ratings on its development set may...
arXiv:2606. 13685v1 Announce Type: cross Abstract: LLM-as-a-Judge is now widely used to rank model outputs, train reward models, and populate public leaderboards, but its run-to-run reliability remains under-characterized.
arXiv:2606. 06546v1 Announce Type: new Abstract: Evaluating large language models (LLMs) for education requires measuring how models teach, not only what they know.
The paper investigates whether consensus among large language model (LLM) judges truly reflects human alignment. By treating each judge’s scores as vectors, the authors measure spread, effective rank, and angles to human scores across 42 judges on Indic benchmarks, revealing that inter‑judge agreement often mirrors shared blind spots rather than human judgments. They find that while judges agree as much as humans, they only reach 58‑66% of human agreement and frequently focus on axes humans do not weight, indicating that ensemble agreement alone is insufficient evidence of alignment.
arXiv:2608.29517v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as essay graders in learning analytics, evaluated almost exclusively with agreement statistics. Educ...