MetaRubric: Learning to Reward for Rubric-Based Reinforcement Learning
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2607. 25659v1 Announce Type: new Abstract: Rubric-based reinforcement learning enriches language model training by evaluating model outputs against explicit criteria.
Long-horizon large language model (LLM) agents are typically optimized with sparse terminal outcomes, making fine-grained credit assignment across multi-step interactions difficult. Existing approache...
arXiv:2608. 16156v1 Announce Type: new Abstract: Long-horizon large language model (LLM) agents are typically optimized with sparse terminal outcomes, making fine-grained credit assignment across multi-step interactions difficult.
arXiv:2606. 03980v1 Announce Type: new Abstract: Reward models (RMs) provide critical feedback signals for LLM post-training, notably in reinforced fine-tuning (RFT) and reinforcement learning (RL) pipelines.
arXiv:2608.30005v1 Announce Type: new Abstract: Rubric-based reinforcement learning extends RL beyond tasks with exact answers or rule-based verifiers by scoring responses against instance-specific c...
The paper identifies a flaw in rubric‑based reinforcement learning where additive reward aggregation allows policies to compensate for missing critical criteria, leading to higher scores but poorer answers, especially in clinical consultation tasks. It demonstrates that grouping rubric criteria into protocol‑level dimensions—so a dimension only counts when all its criteria are satisfied—mitigates this reward hacking. The proposed Protocol‑level Rubrics (ProRubric) improve appropriateness by 10.8 points without sacrificing coverage and achieve the best performance across seven benchmarks.