The paper introduces ImpossibleRubrics, a benchmark of 169 impossible tasks designed to test the robustness of language‑model‑generated rubrics as reward signals. Each task is paired with a verifiable oracle certificate that defines what constitutes an honest answer, and the benchmark includes 48 answerable controls. Experiments show that many rubric generators are exploited frequently—up to 36% on a stress cut—highlighting a significant gap in rubric quality rather than task difficulty, and that generic rubrics can be more vulnerable than tailored ones.
By Bowen Qin, Yi Xie, Yesheng Liu, Xi Yang
arXiv:2608.30005v1 Announce Type: new
Abstract: Rubric-based reinforcement learning extends RL beyond tasks with exact answers or rule-based verifiers by scoring responses against instance-specific c...
By Fengyu Xie, Yilun Zhao, Bingsen Chen, Arman Cohan, Chen Zhao
arXiv:2609.23457v1 Announce Type: new
Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) is expanding from tasks with well-defined correctness signals, such as mathematics and code, towa...
By Hao Li, Zhengkun Zhang, Gangqiang Hu, Zhen Zhang, Yude Gao, Dai Dai, Jing Liu
arXiv:2603. 00077v3 Announce Type: replace-cross Abstract: Rubric-based LLM judges have become indispensable for evaluating and optimizing systems on non-verifiable tasks, where success cannot be reduced to exact programmatic checks.
By Delip Rao, Chris Callison-Burch
arXiv:2609.01354v1 Announce Type: cross
Abstract: Reinforcement learning with verifiable rewards (RLVR) and standard benchmark evaluation both rely on an automatic verifier that turns a free text ans...
By Esther Xin
Reinforcement Learning with Verifiable Rewards (RLVR) is expanding from tasks with well-defined correctness signals, such as mathematics and code, toward multifaceted quality requirements specified by...
The paper introduces AdaptRubric, a Coarse-to-Fine Rubrics Framework designed to create task‑adaptive judging criteria for GUI reward modeling. It first retrieves a category‑level coarse rubric by mapping instructions to a GUI task family, then generates an instance‑level fine rubric that captures specific values, scopes, and constraints from the instruction. Experiments show that AdaptRubric outperforms existing reward agents, improving F1 by 3.6 points and achieving a 4.23‑point task‑success gain under a matched image budget.
By Tao Xiong, Xavier Hu, Wenkai Wang, Qinzhuo Wu, Changqiao Wu, Pengzhi Gao, Wei Liu, Jian Luan, Shengyu Zhang
arXiv:2605.14040v2 Announce Type: replace
Abstract: Trackable improvement in multimodal physics reasoning rests on a training-and-evaluation system that is itself rarely verified: the corpora a model...
By Shan Yang
The paper introduces a penalty‑aware evaluation framework for Retrieval‑Augmented Generation (RAG) systems that uses asymmetric scoring, knowledge‑gap canaries, and a failure‑attribution pipeline. Applying this framework to three commercial RAG products and a baseline on SimpleQA‑Verified, the authors find that while overall accuracy is similar across systems, canary violation rates vary dramatically, showing that systems differ more in when they answer than in what they answer. The study demonstrates that penalty‑aware scoring can reorder system rankings and is robust across different penalty settings.
By Alden Do Rosario, Hussein Younes, Felipe Pires
arXiv:2606. 26300v1 Announce Type: new Abstract: A classical intuition holds that verifying a solution is easier than producing one.
By Binghai Wang, Chenlong Zhang, Dayiheng Liu, Jiajun Zhang, Jiawei Chen, Mouxiang Chen, Rongyao Fang, Siyuan Zhang, Xuwu Wang, Yuheng Jing, Zeyao Ma, Zeyu Cui
arXiv:2608. 11669v1 Announce Type: cross Abstract: Reinforcement learning against rubrics, lists of criteria graded by an LLM judge, has become a standard way to post-train language models on tasks with no deterministic answer.
By Minglai Yang, Xinyu Guo, Utkarsh Tyagi, Mian Zhang, Razvan Dumitru, Sunjie Hou, Yunzhong He, Daniel Yue Zhang, Ying Liu
arXiv:2510. 00492v3 Announce Type: replace Abstract: The reliability of large language models (LLMs) during test-time scaling is often assessed with \emph{external verifiers} or \emph{reward models} that distinguish correct reasoning from flawed logic.
By Dong Bok Lee, Seanie Lee, Sangwoo Park, Minki Kang, Jinheon Baek, Dongki Kim, Dominik Wagner, Jiongdao Jin, Heejun Lee, Tobias Bocklet, Jinyu Wang, Jingjing Fu, Sung Ju Hwang, Jiang Bian, Lei Song