The article examines how classical test theory reliability statistics misrepresent the performance of large language model (LLM) judges. It shows that internal‑consistency coefficients, the dependability index, and Livingston‑Lewis accuracy each conflate judge error with item design or criterion validity, making it impossible to attribute a single reliability value to the judge alone. The authors argue that such misattribution can influence deployment decisions and documentation.
By Louis Yiven Zhu
arXiv:2609.07162v1 Announce Type: new
Abstract: Several properties safety monitors are asked to certify, among them cross-tenant noninterference, sandbagging and evaluation awareness, are 2-safety hy...
By Xin Xu
The paper investigates how language‑model judges can make version‑dependent errors when evaluating upgraded agents. Using 35 public coding‑agent submissions, two customer‑service agents, and over a thousand expert‑labeled trajectories, the authors show that fixed judges often reject task‑conditioned error invariance and can incorrectly approve failed patches, especially as agent capability increases. Paired audits of current outputs reduce interval width only marginally, and the study concludes that independent human patch review is still necessary.
By Jiapeng Li
The paper investigates how many human annotators are equivalent to a panel of 32 large‑language‑model (LLM) judges. By comparing the panel’s label distributions to empirical human labels on three ChaosNLI tasks, the authors find two distinct effective panel sizes: distribution‑error matching yields effective sizes of 2.304, 3.750, and 3.445, while spectral matching gives 4.242, 6.459, and 6.499, indicating a 1.72–1.89× gap. The study also explores how spectral diversity, participation ratio, and panel composition affect effective size, and demonstrates that carefully chosen panels can outperform baseline accuracy while improving effective size.
By Chao Li, Yingying Yu, Yunfeng Li
Enterprise practitioners read agent leaderboards as if they ranked agent capability. We show, across three open agent-trace benchmarks (TheAgentCompany, $τ^2$-bench, and AppWorld), that the agent main effect accounts for less than 3% of total variance in every dataset and check type, while the agent-by-task interaction accounts for 7-23%.
The study analyzes 373,019 judgments from LLM‑scored benchmarks, decomposing variance into system, item, judge, and interaction components via generalizability theory. It finds that with a single judge, generalizability converges to a ceiling determined by the system‑by‑judge variance, which is substantially lower in pairwise preference settings, allowing one judge to suffice. The research also reveals significant biases in presentation order and highlights that many published win‑rate claims fall below the measured floor of the benchmarks.
By Atul Anand
The article examines how classical test theory statistics—Kuder‑Richardson coefficient, dependability index, and Livingston‑Lewis accuracy—can mislead when applied to large language model (LLM) judges that are evaluated with a single prompt and no gold labels. Using Claude Haiku 4.5 on 210 short‑answer items, the authors show that these metrics fail to isolate the judge’s performance because the judge’s single administration provides no variance component. They argue that reliable statements about an LLM judge require gold labels or varied scorer facets, and that bank design heavily influences reliability estimates.
By Louis Yiven Zhu
arXiv:2605. 27914v2 Announce Type: replace-cross Abstract: Benchmarking is mature where answers are verifiable -- math, code, reasoning -- but the fastest-growing uses of LLMs are subjective and human-facing: companionship, emotional support, counseling.
By Yuming (Rapheal), Huang, Yao Liu, Pengjie Ding, Lei Wang, Junchen Wan
The paper demonstrates that a preference‑optimization objective can learn to distinguish reliable from unreliable sources by installing a prior‑dependent reliability switch. By training on data where a source’s stated reliability is paired with its answer, the model learns to flip its response only when the stated reliability exceeds a threshold that grows with the model’s prior. Experiments on Qwen2.5‑7B‑Instruct and Llama‑3.1‑8B show that this switch generalizes to unseen reliability values and follows stated reliability over role prestige, whereas supervised imitation fails to learn it.
By Sen Yang, Yuen-Hei Yeung
arXiv:2608. 11323v1 Announce Type: new Abstract: Enterprise practitioners read agent leaderboards as if they ranked agent capability.
By Vasundra Srinivasan
The paper examines safety routers—systems that route user requests to different language models—and finds that their performance degrades significantly when evaluated under distribution shift. In standard benchmarks, routers appear effective because the best single model is chosen from the same evaluation data, but when the data distribution changes, the routing advantage diminishes or disappears. The study quantifies this bias across multiple safety corpora, showing that routers offer little benefit under realistic shift conditions and that recognition‑based defenses can be undermined by attackers who know the model being used.
By Amit Singh Bhatti, Vishal Vaddina
The paper presents a closed‑form estimator for the contamination correlation between anchors and judges under a single‑common‑factor model, requiring at least two judges and two anchors. It introduces a diagnostic battery—including judge‑covariance dispersion, over‑identification tests, a family‑block test, bootstrap confidence intervals, and a weak‑identification screen—to validate the estimator and detect violations. The authors also discuss identification limits for ordinal data and report that real panels have not yet passed the model‑adequacy pre‑test, while simulation studies confirm the estimator’s performance.
By Veerendra Kumar Sunkavalli