The paper investigates when a language‑model judge can truly ground its verdicts in code correctness. It shows that current multi‑agent verification methods rely on evidence that is both independent of the answer and distinct between candidates—conditions that fail in code judging. By analyzing two label‑free measurements from the judge’s logs, the authors demonstrate that gating on one measurement allows the system to decline uncertain comparisons, improving accuracy from 20.7% to 36.9% while still answering half of all cases.
By Salma Roshdy Aly, Hussein Assaf, Ziad Kobti
arXiv:2601.08654v3 Announce Type: replace
Abstract: Rubric-based text evaluation increasingly relies on large language models (LLMs) as scalable judges, yet frozen black-box models can interpret the...
By Yihan Hong, Huaiyuan Yao, Bolin Shen, Wanpeng Xu, Hua Wei, Yushun Dong
arXiv:2609.13824v1 Announce Type: new
Abstract: Large language models (LLMs) are increasingly used to evaluate the responses of other language models. This approach, known as LLM-as-a-Judge, is faste...
By Aakash Kumar Tiwari
arXiv:2609.26550v3 Announce Type: replace
Abstract: LLM-as-a-judge scales evaluation, but reasoning judges are slow and costly. We study JEV-as-a-Judge: evaluation with JEV, a decision-only judge tha...
By Yubo Li, Yidi Miao, Ramayya Krishnan, Rema Padman
The paper investigates whether consensus among large language model (LLM) judges truly reflects human alignment. By treating each judge’s scores as vectors, the authors measure spread, effective rank, and angles to human scores across 42 judges on Indic benchmarks, revealing that inter‑judge agreement often mirrors shared blind spots rather than human judgments. They find that while judges agree as much as humans, they only reach 58‑66% of human agreement and frequently focus on axes humans do not weight, indicating that ensemble agreement alone is insufficient evidence of alignment.
By Sourabrata Mukherjee, Hamna Hamna, Kalika Bali, Sunayana Sitaram
The paper introduces a two‑dimensional construct validity framework for evaluating large language models (LLMs) as judges, defining invariance (S) and sensitivity (R) to construct‑preserving and construct‑changing edits. Experiments across seven judges and four domains reveal high invariance (average S = 0.945) but low sensitivity (average R = 0.319), with sensitivity varying by edit type. Audits of public label sets show that surface‑only predictors can reproduce a substantial portion of labels, underscoring that high agreement does not guarantee construct validity.
By Jianlin Chen, Wenhui Chen, Ziyao Lin, Chi Man Vong