MatrixReward: Reward from Rubric Matrix for Open-Ended Generation
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
Reinforcement Learning with Verifiable Rewards (RLVR) is expanding from tasks with well-defined correctness signals, such as mathematics and code, toward multifaceted quality requirements specified by...
arXiv:2609.23457v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) is expanding from tasks with well-defined correctness signals, such as mathematics and code, towa...
The paper introduces a method for creating query‑specific rubrics for DeepResearch‑style long‑form report generation by training rubric generators with reinforcement learning. It builds a dataset of queries annotated with human preferences, then uses a hybrid reward that includes preference consistency, format validity, and LLM‑based rubric evaluation. The learned rubrics outperform generic or manually constructed alternatives in distinguishing preferred reports and, when used as rewards, improve performance of both single‑agent and multi‑agent DeepResearch systems.
arXiv:2609.36652v1 Announce Type: new Abstract: Open-ended generation lacks canonical answers, making pointwise rewards difficult to calibrate for group-based reinforcement learning. Directly ranking...
arXiv:2609.22947v1 Announce Type: new Abstract: Reinforcement learning (RL) is vital for optimizing video generation models, with a robust reward model (RM) serving as the cornerstone. However, exist...
arXiv:2605.26958v2 Announce Type: replace-cross Abstract: Reinforcement learning in open-ended long-form generation is challenging because reliable reference answers and automatic metrics are often u...