The paper introduces ImpossibleRubrics, a benchmark of 169 impossible tasks designed to test the robustness of language‑model‑generated rubrics as reward signals. Each task is paired with a verifiable oracle certificate that defines what constitutes an honest answer, and the benchmark includes 48 answerable controls. Experiments show that many rubric generators are exploited frequently—up to 36% on a stress cut—highlighting a significant gap in rubric quality rather than task difficulty, and that generic rubrics can be more vulnerable than tailored ones.
By Bowen Qin, Yi Xie, Yesheng Liu, Xi Yang
arXiv:2608.30005v1 Announce Type: new
Abstract: Rubric-based reinforcement learning extends RL beyond tasks with exact answers or rule-based verifiers by scoring responses against instance-specific c...
By Fengyu Xie, Yilun Zhao, Bingsen Chen, Arman Cohan, Chen Zhao
arXiv:2609.23457v1 Announce Type: new
Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) is expanding from tasks with well-defined correctness signals, such as mathematics and code, towa...
By Hao Li, Zhengkun Zhang, Gangqiang Hu, Zhen Zhang, Yude Gao, Dai Dai, Jing Liu
arXiv:2603. 00077v3 Announce Type: replace-cross Abstract: Rubric-based LLM judges have become indispensable for evaluating and optimizing systems on non-verifiable tasks, where success cannot be reduced to exact programmatic checks.
By Delip Rao, Chris Callison-Burch
arXiv:2609.01354v1 Announce Type: cross
Abstract: Reinforcement learning with verifiable rewards (RLVR) and standard benchmark evaluation both rely on an automatic verifier that turns a free text ans...
By Esther Xin
Reinforcement Learning with Verifiable Rewards (RLVR) is expanding from tasks with well-defined correctness signals, such as mathematics and code, toward multifaceted quality requirements specified by...