arXiv AI By Wisdom Ikezogwo, Mehmet Saygin Seyfioglu, Ranjay Krishna, Karim Bouyarmane

When Rubrics Fail: Error Enumeration as Reward in Reference-Free RL Post-Training for Virtual Try-On

Read the original on arXiv AI →

arXiv:2603. 05659v3 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards (RLVR) and Rubrics as Rewards (RaR) have driven strong gains in domains with clear correctness signals and even in subjective domains by synthesizing evaluation criteria from ideal reference answers.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 16

ImpossibleRubrics: Stress-Testing Generated Rubrics as Reward Signals

The paper introduces ImpossibleRubrics, a benchmark of 169 impossible tasks designed to test the robustness of language‑model‑generated rubrics as reward signals. Each task is paired with a verifiable oracle certificate that defines what constitutes an honest answer, and the benchmark includes 48 answerable controls. Experiments show that many rubric generators are exploited frequently—up to 36% on a stress cut—highlighting a significant gap in rubric quality rather than task difficulty, and that generic rubrics can be more vulnerable than tailored ones.

By Bowen Qin, Yi Xie, Yesheng Liu, Xi Yang