arXiv AI

Scoring Higher, Answering Worse: Mitigating Reward Hacking in Rubric-Based RL via Protocol-Level Rubrics

The paper identifies a flaw in rubric‑based reinforcement learning where additive reward aggregation allows policies to compensate for missing critical criteria, leading to higher scores but poorer answers, especially in clinical consultation tasks. It demonstrates that grouping rubric criteria into protocol‑level dimensions—so a dimension only counts when all its criteria are satisfied—mitigates this reward hacking. The proposed Protocol‑level Rubrics (ProRubric) improve appropriateness by 10.8 points without sacrificing coverage and achieve the best performance across seven benchmarks.

arXiv AI
Aug 13

Rubric Dropout: A Simple Way to Mitigate Reward Hacking in Rubric-as-Reward RL

arXiv:2608. 11669v1 Announce Type: cross Abstract: Reinforcement learning against rubrics, lists of criteria graded by an LLM judge, has become a standard way to post-train language models on tasks with no deterministic answer.

By Minglai Yang, Xinyu Guo, Utkarsh Tyagi, Mian Zhang, Razvan Dumitru, Sunjie Hou, Yunzhong He, Daniel Yue Zhang, Ying Liu
arXiv AI
Aug 5

Rubrics as Privileged Information for Open-Ended Generation

arXiv:2608. 02948v1 Announce Type: cross Abstract: On-policy self-distillation (OPSD), where a single model acts as both student and teacher with different contexts, has shown promise in verifiable domains like math, where hard privileged information (PI) in the form of ground-truth answers structurally constrains valid continuations.

By Deepika Bablani, Ajay Gupta, Wanming Chen
arXiv AI
Aug 17

Inducing Reward-Free Judging Rubrics that Reduce Over-Crediting in Agent Evaluation

arXiv:2608. 13564v1 Announce Type: new Abstract: Evaluating language-model agents at scale increasingly relies on a second language model as an automatic judge, because the gold signal, an executable environment reward, is expensive, slow, or unavailable at deployment time.

By Darragh Quinn, David Dylan, Roisin Healy, Fionn Carroll, Maeve Donnelly, Cormac Sheehan
arXiv Computation and Language
4d ago

AdaTutoRank: Learning to Rerank Document Sets via Adaptive Tutoring Optimization for RAG and Deep Research

AdaTutoRank introduces a setwise document reranker that uses Adaptive Tutoring Optimization (ATO) to provide graded supervision across nine rubric dimensions. By generating hint‑based silver labels, reinforcement rewards, and distillation cues tailored to each rollout’s quality, the method improves credit assignment for individual documents within a set. Experiments on ten benchmarks covering Retrieval‑Augmented Generation (RAG), deep research, and setwise evaluation show that AdaTutoRank achieves superior overall performance while reducing the number of retrieval calls.

By Kailin Jiang, Lei Liu, Jian Xi, Yangqi Chen, Hui Xu, Hongwei Zhao, Bin Li, Yu Lu, Haibo Shi
arXiv AI
Jul 7

When Rubrics Fail: Error Enumeration as Reward in Reference-Free RL Post-Training for Virtual Try-On

arXiv:2603. 05659v3 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards (RLVR) and Rubrics as Rewards (RaR) have driven strong gains in domains with clear correctness signals and even in subjective domains by synthesizing evaluation criteria from ideal reference answers.

By Wisdom Ikezogwo, Mehmet Saygin Seyfioglu, Ranjay Krishna, Karim Bouyarmane