Metric Match: A Subset Selection Approach to Evaluating LLM Judge Reliability
arXiv:2606. 15029v1 Announce Type: new Abstract: LLM judges are used to reduce the need for costly human labor in evaluating open-ended text generation.
arXiv:2603. 10400v2 Announce Type: replace-cross Abstract: Designing service systems requires selecting among alternative configurations -- choosing the best chatbot variant, the optimal routing policy, or the most effective quality control procedure.
arXiv:2606. 15029v1 Announce Type: new Abstract: LLM judges are used to reduce the need for costly human labor in evaluating open-ended text generation.
arXiv:2609.38860v1 Announce Type: cross Abstract: Learning from human preferences is central to large language model (LLM) alignment, but human preference annotation is costly. Active preference lear...
arXiv:2609.33676v2 Announce Type: replace Abstract: LLM agents increasingly take consequential actions through interactions with users, policies, and external tools. Auditing these agents requires au...
arXiv:2608. 13840v1 Announce Type: cross Abstract: Audits of generative AI (GenAI) systems often summarize behavior as a reported rate: how often the audited system complies with policy.
The paper introduces a framework for evaluating AI systems that not only checks final labels but also tracks the reasoning behind them through three core sources—grounds, norms, and authority—forming an eight-cell counterfactual judgment cube. It defines minimal source replacement sets, called judgment receipts, to explain changes in verdicts and provides certification cost bounds for black-box evaluators. The authors present ReasonBench, a benchmark with 19,520 cases, and demonstrate that while high standard accuracy can mask robustness issues, receipt accuracy reveals significant gaps in reasoning consistency across different models.
arXiv:2609.39026v1 Announce Type: new Abstract: Deep Research agents synthesize evidence into cited reports, yet a well-cited report can still reach a misleading conclusion. Citation correctness chec...
arXiv:2606. 30338v1 Announce Type: new Abstract: External evaluations are becoming increasingly central to the governance of AI systems.
The paper introduces a framework for hypothesis testing that combines inexpensive AI judgments with selective human verification to control type‑I and type‑II errors while minimizing cost. It derives an information‑theoretic lower bound on the minimum cost and proposes the SCALE policy, a sequential, cost‑aware strategy that adapts AI scoring and human escalation. SCALE is proven valid for finite samples and asymptotically matches the lower bound, achieving significant savings when both AI and human inputs are valuable.
The paper introduces prediction‑powered smoothing (PP‑S) and its taxonomy‑aware extension (PP‑TS) to improve point and interval estimates of domain‑specific AI performance when only a limited sample of labeled units is available. It also proposes a new design‑based cross‑validation score that is approximately unbiased for selecting between direct and smoothed estimators. Experiments on a curated benchmark and real‑world agent traffic show that the proposed methods outperform direct estimators in both accuracy and coverage, and that the new score matches the performance of an independent validation sample while providing more precise error estimates.
arXiv:2608. 08700v1 Announce Type: new Abstract: Reliable evaluation of tool routing is critical as Large Language Models increasingly operate as autonomous agents.
Reference-based text evaluation metrics, which are widely used to assess natural language generation systems, score a candidate response by comparing it with a reference response. The reliability of an evaluation metric is usually judged by its statistical correlation with human ratings.
The paper introduces a validation‑gated audit framework for voice AI customer‑care systems, treating them as stateful, multi‑turn, tool‑mediated interactions where bias and safety can manifest as added burdens before a final decision. The framework distinguishes between native speech‑to‑speech, cascaded ASR‑to‑LM‑to‑TTS, and hybrid architectures, and applies matched service facts across controlled caller presentation conditions to validate fact invariance, presentation cues, artifacts, and acoustic measurements. It outlines seven validation gates, a six‑family metric set, and demonstrates the approach with a synthetic refund‑dispute audit example, while noting that production results are withheld until the protocol is satisfied.