arXiv AI By Tzu-Heng Huang, Shengqi Qiu, Frederic Sala

Codifying the Judge: Scalable Evaluation via Program Distillation

Read the original on arXiv AI →

arXiv:2607. 22561v1 Announce Type: new Abstract: LLM-as-a-judge has become the standard for automated evaluation, but it suffers from high cost, significant latency, and opaque decisions -- limitations that undermine its scalability and reliability.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 14

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces

arXiv:2608. 12585v1 Announce Type: new Abstract: Improving reasoning LLMs requires the ability to judge the quality of long reasoning traces for effective reasoning data curation, strong training signals during reinforcement learning, and an in-depth understanding of reasoning behaviors during model performance evaluation.

By Congchao Wang, Diwakar Singh, Qiaozi Gao, Spyros Matsoukas, Yang Liu, Mahdi Namazifar
arXiv AI
4d ago

From Judgment Quality to Downstream Utility: Rethinking LLM-as-a-Judge for Open-Ended Tasks

The paper investigates how the design of LLM-as-a-Judge protocols influences both the intrinsic quality of judgments and their downstream utility in open-ended tasks. By varying verdict granularity, critique usage, and evaluation batching, and by applying Judge guidance to test-time inference methods such as Best-of-N selection, revision, and beam search, the authors find that judgment quality and downstream performance do not always align and that protocol choices significantly affect outcomes. The study highlights the need for comprehensive evaluation of LLM Judges that considers both judgment quality and practical utility.

By Zheng Zhang, Lufei Li, Xinyue Tan, Yuanhao Zeng, Ziwei Shan, Yexin Li, Kan Ren
arXiv Computation and Language
Aug 28

JudgeStealer: Extracting LLM Judging Capabilities across Evaluation Protocols

JudgeStealer is a query‑efficient framework that extracts the judging capabilities of large language models across pointwise scoring, pairwise comparison, and listwise ranking protocols. It leverages cross‑protocol agreement to convert pointwise scores into higher‑order supervision, dynamically selects informative inputs, and applies score smoothing and multi‑protocol review to preserve ordinal structure and avoid catastrophic forgetting. Experiments show it outperforms existing baselines, achieving up to 73.3% accuracy on pointwise, 87.0% on pairwise, and 71.6% on listwise evaluation, while remaining robust against common extraction defenses.

By Chen Chen, Yaolin Chen, Xuehan Sun, Juan Lin, Xueluan Gong, Yuhang Zheng, Qian Wang, Kwok-Yan Lam