arXiv Machine Learning By Can Lv, Mingju Chen, Heng Chang, Shiji Zhou

Mitigating False Credit Propagation: Probabilistic Graphical Reward Aggregation for Rubric-Based Reinforcement Learning

Read the original on arXiv Machine Learning →

arXiv:2606. 03361v1 Announce Type: new Abstract: Rubric-based rewards are increasingly used for open-ended language model post-training, but criterion-level scores are often aggregated as independent utilities.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 4

DRACO: Fine-Grained Credit Assignment with Dynamic Rubrics for Long-Horizon Agent Training

DRACO introduces a method for fine‑grained credit assignment in long‑horizon reinforcement learning tasks that lack verifiable rewards. It dynamically generates multi‑criteria rubrics during training, scores them once per trajectory, and redistributes the resulting judgment over the steps responsible for each rubric to produce differentiated per‑step advantages. Experiments on AppWorld and Tau‑Bench show significant performance gains over baseline models and other rubric‑based approaches, even without using any verifiers.

By Shubham Gandhi, Saurabh Goyal, Kiran Kate, Yara Rizk
Hugging Face Trending Papers
Sep 3

DRACO: Fine-Grained Credit Assignment with Dynamic Rubrics for Long-Horizon Agent Training

DRACO introduces a method for fine‑grained credit assignment in long‑horizon reinforcement learning tasks that lack verifiable rewards. It dynamically generates multi‑criteria rubrics during training, scores them once per trajectory, and redistributes the resulting judgment over the steps responsible for each rubric to produce differentiated per‑step advantages. Experiments on AppWorld and Tau‑Bench show that DRACO outperforms baseline models and other rubric‑based approaches, achieving significant performance gains without relying on verifiers.

arXiv AI
Aug 17

Inducing Reward-Free Judging Rubrics that Reduce Over-Crediting in Agent Evaluation

arXiv:2608. 13564v1 Announce Type: new Abstract: Evaluating language-model agents at scale increasingly relies on a second language model as an automatic judge, because the gold signal, an executable environment reward, is expensive, slow, or unavailable at deployment time.

By Darragh Quinn, David Dylan, Roisin Healy, Fionn Carroll, Maeve Donnelly, Cormac Sheehan
arXiv AI
2d ago

Dependency-Aware Reward Shaping for Agentic Reinforcement Learning

The paper introduces Dependency‑Aware Reward Shaping (DARS), a method that assigns step‑level credit in reinforcement learning by modeling task progress as a graph of predicates with prerequisite relations. Annotators mark each step’s effect on predicates, and DARS discounts verified predicates based on distance from broken prerequisites while preserving independent ones, converting these annotations into signed per‑step rewards. Experiments on five task families with models ranging from 1.5B to 8B show that DARS improves success rates by up to 10 points over GiGPO, boosts WebShop and Search‑R1 QA scores, complements AEPO on AIME24/25, and outperforms OmniOPD in tool‑free reasoning, with ablations confirming the contribution of step‑level credit, dependency attenuation, and graph topology.

By Ziyi Chen, Yan Zhang, Jianhui Wei, Daoan Zhang, Zuozhu Liu
arXiv AI
Aug 13

Rubric Dropout: A Simple Way to Mitigate Reward Hacking in Rubric-as-Reward RL

arXiv:2608. 11669v1 Announce Type: cross Abstract: Reinforcement learning against rubrics, lists of criteria graded by an LLM judge, has become a standard way to post-train language models on tasks with no deterministic answer.

By Minglai Yang, Xinyu Guo, Utkarsh Tyagi, Mian Zhang, Razvan Dumitru, Sunjie Hou, Yunzhong He, Daniel Yue Zhang, Ying Liu