arXiv AI

Backbone-Adaptive Evidence Routing for Robust Pairwise LLM Judging

Backbone-Adaptive Evidence Routing (BAER) is a method that dynamically selects an evidence gathering protocol—evidence stacking, reliability-based expert routing, or candidate-blind reference verification—based on the judge backbone and benchmark during development, then locks that choice for testing. BAER maintains candidate symmetry, ensuring that swapping responses reverses preference but not its strength. In experiments across four benchmarks and two 8B judge backbones, BAER outperforms all compared methods, achieving the highest test accuracy in all eight conditions and improving over the strongest baseline by 0.87–7.32 points.

arXiv AI
Aug 24

JuryProbe: An Empirical Consensus-Risk Diagnostic for Routing Reference-Free Factuality Judge Panels to Grounded Verification

JuryProbe is an empirical diagnostic tool designed to assess consensus risk in panels of reference‑free large language model judges used for factuality verification. It estimates risk by measuring false‑negative correlations and false‑consensus lift from a labeled calibration probe, and routes high‑risk majority decisions to judges with trusted references. The approach was validated on FEVER corruptions, showing that flagged decisions can be grounded without additional reference acquisition in most cases, while reducing false accepts by about 0.4% and avoiding 28% of reference acquisitions.

By Tianxin Zhou, Ruixi Lin
arXiv Computation and Language
Aug 28

JudgeStealer: Extracting LLM Judging Capabilities across Evaluation Protocols

JudgeStealer is a query‑efficient framework that extracts the judging capabilities of large language models across pointwise scoring, pairwise comparison, and listwise ranking protocols. It leverages cross‑protocol agreement to convert pointwise scores into higher‑order supervision, dynamically selects informative inputs, and applies score smoothing and multi‑protocol review to preserve ordinal structure and avoid catastrophic forgetting. Experiments show it outperforms existing baselines, achieving up to 73.3% accuracy on pointwise, 87.0% on pairwise, and 71.6% on listwise evaluation, while remaining robust against common extraction defenses.

By Chen Chen, Yaolin Chen, Xuehan Sun, Juan Lin, Xueluan Gong, Yuhang Zheng, Qian Wang, Kwok-Yan Lam
Hugging Face Trending Papers
Aug 18

Judge, Retrieve, or Abstain: Uncertainty-Guarded LLM Judging with Provable Risk Guarantees

The paper introduces a risk‑controlled framework for using large language models (LLMs) as judges in tasks without reference answers. By calibrating uncertainty thresholds on a held‑out set, the method ensures that the false discovery rate of accepted verdicts stays below a user‑specified level α with high probability, using finite‑sample Clopper–Pearson intervals. When the parametric judge lacks confidence, the instance is routed to a retrieval‑augmented mode with a second calibrated threshold, preserving the error guarantee while achieving higher coverage than single‑mode baselines.

arXiv Computation and Language
Aug 31

Benchmarking LLM-as-a-Judge for Long-Form Output Evaluation

The paper introduces LongJudgeBench, a benchmark designed to evaluate large language models (LLMs) acting as judges for long-form text generation. It highlights that long-form evaluation requires complex, document-level assessments beyond simple length, such as organization, coverage, depth, consistency, and scenario-specific quality. Experiments show a significant reliability gap among current LLM judges, indicating instability across scenarios and limited effectiveness of rubrics or references.

By Junjie Chen, Yuxi Dong, Haitao Li, Weihang Su, Yujia Zhou, Min Zhang, Yiqun Liu, Qingyao Ai
arXiv Machine Learning
Sep 14

Are Independently Estimated View Uncertainties Comparable? Unified Routing for Trusted Multi-View Classification

The paper introduces Trusted Multi-view learning with Unified Routing (TMUR), a method that separates view-specific evidence extraction from fusion arbitration in multi-view classification. TMUR employs view-private experts, a collaborative expert, and a unified router that assigns sample-level weights based on global context, along with soft load-balancing and diversity regularization to promote balanced and discriminative expert use. Experiments on 14 datasets show that TMUR consistently improves classification accuracy and reliability compared to 15 recent baselines.

By Yilin Zhang, Cai Xu, Haishun Chen, Ziyu Guan, Wei Zhao
arXiv AI
Aug 24

No Judgment Without a Reason: Counterfactual Receipts for Versioned AI Evaluators

The paper introduces a framework for evaluating AI systems that not only checks final labels but also tracks the reasoning behind them through three core sources—grounds, norms, and authority—forming an eight-cell counterfactual judgment cube. It defines minimal source replacement sets, called judgment receipts, to explain changes in verdicts and provides certification cost bounds for black-box evaluators. The authors present ReasonBench, a benchmark with 19,520 cases, and demonstrate that while high standard accuracy can mask robustness issues, receipt accuracy reveals significant gaps in reasoning consistency across different models.

By Ye Chen, Weining Zhang
arXiv Machine Learning
Sep 14

Can We Trust LLM Judges: A Study of Capability-Dependent Biases and Multi-Judge Ensemble for Bias Calibration

The paper investigates how large language models (LLMs) used as judges in absolute scoring tasks exhibit systematic biases that compromise reliability. It shows that a judge’s task accuracy strongly predicts both its judging accuracy and its directional bias, yet more capable examinee models consistently receive more lenient judgments. To mitigate these biases, the authors propose a calibrated weighted majority voting (WMV) ensemble that estimates judges’ error rates from inter-judge agreement patterns, achieving near-oracle performance without labeled data and improving both accuracy and fairness.

By Gemma Zhang, Prachi Badarayani, Asmi Kumar, Sadid Hasan, Sulaiman Vesal
arXiv AI
Sep 24

COMED: The Missing Middle Between Routing and Collaboration in Multi-LLM Inference

COMED (Controlled Model Escalation for Multi-LLM Deliberation) is a post-anchor controller that selectively engages cross-model collaboration in multi-LLM inference. It uses anchor self‑consistency, router margin, and a lightweight peer probe to accept confident answers, verify ambiguous cases, and only escalates when collaboration is likely beneficial. Experiments on medical, scientific, and general reasoning benchmarks show that COMED improves performance across 16 open‑weight settings, achieving up to +10.7 percentage points on MedQA and outperforming dense collaboration while invoking fewer models and decoded tokens.

By Norah Alballa, Wenxuan Zhang, Salma Kharrat, Fares Fourati, Zafar Ayyub Qazi, Mohamed Elhoseiny, Marco Canini