arXiv AI By Darragh Quinn, David Dylan, Roisin Healy, Fionn Carroll, Maeve Donnelly, Cormac Sheehan

Inducing Reward-Free Judging Rubrics that Reduce Over-Crediting in Agent Evaluation

Read the original on arXiv AI →

arXiv:2608. 13564v1 Announce Type: new Abstract: Evaluating language-model agents at scale increasingly relies on a second language model as an automatic judge, because the gold signal, an executable environment reward, is expensive, slow, or unavailable at deployment time.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 13

Rubric Dropout: A Simple Way to Mitigate Reward Hacking in Rubric-as-Reward RL

arXiv:2608. 11669v1 Announce Type: cross Abstract: Reinforcement learning against rubrics, lists of criteria graded by an LLM judge, has become a standard way to post-train language models on tasks with no deterministic answer.

By Minglai Yang, Xinyu Guo, Utkarsh Tyagi, Mian Zhang, Razvan Dumitru, Sunjie Hou, Yunzhong He, Daniel Yue Zhang, Ying Liu
arXiv AI
Sep 2

Commit-first LLM judging inherits the judge's own errors

The paper investigates whether widely used evaluation frameworks for large language models (LLMs) implement a defense called commit‑first judging, which requires a judge to solve a task itself before accepting a candidate answer. Across 24 configurations in eight popular frameworks, none use the full commit‑first method; nine use a weaker variant that is ineffective. In controlled experiments, the weaker variant allowed systems to game the judge, while the full commit‑first approach eliminated this vulnerability but sometimes worsened evaluation when the judge’s own answer was incorrect.

By Idil Gozel
Hugging Face Trending Papers
Sep 3

DRACO: Fine-Grained Credit Assignment with Dynamic Rubrics for Long-Horizon Agent Training

DRACO introduces a method for fine‑grained credit assignment in long‑horizon reinforcement learning tasks that lack verifiable rewards. It dynamically generates multi‑criteria rubrics during training, scores them once per trajectory, and redistributes the resulting judgment over the steps responsible for each rubric to produce differentiated per‑step advantages. Experiments on AppWorld and Tau‑Bench show that DRACO outperforms baseline models and other rubric‑based approaches, achieving significant performance gains without relying on verifiers.