Which Constraints Are Missing? Ask the Verifier: Graded Rewards for Constraint-Following Music Generation
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2603. 05659v3 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards (RLVR) and Rubrics as Rewards (RaR) have driven strong gains in domains with clear correctness signals and even in subjective domains by synthesizing evaluation criteria from ideal reference answers.
arXiv:2607. 05904v1 Announce Type: new Abstract: Training a language model against its own reference-free judgments (the premise of self-rewarding, self-play, and LLM-as-a-judge pipelines) assumes a model's verdict on a shown answer tracks correctness.
The paper introduces Reinforcement Learning with Verifiable Rewards (RLVR) applied to small search agents, specifically training a Qwen3.5-0.8B model with Group Relative Policy Optimization and an interleaved Wikipedia-search tool on the MuSiQue dataset. Experiments varying reward shapes across three seeds show that RLVR can achieve a 3.8‑fold improvement over an untrained baseline, with the best run reaching a 0.352 average exact match. The study finds that the sparse exact‑match reward, standard in larger models, performs poorly for small models, indicating that reward design must be tailored rather than scaled down from large‑model recipes.
DRACO introduces a method for fine‑grained credit assignment in long‑horizon reinforcement learning tasks that lack verifiable rewards. It dynamically generates multi‑criteria rubrics during training, scores them once per trajectory, and redistributes the resulting judgment over the steps responsible for each rubric to produce differentiated per‑step advantages. Experiments on AppWorld and Tau‑Bench show that DRACO outperforms baseline models and other rubric‑based approaches, achieving significant performance gains without relying on verifiers.
The paper introduces Circuit Reasoning Score (CRS), a data‑selection signal for reinforcement learning with verifiable rewards that uses attention‑head activity from a frozen base model to gauge reasoning engagement. CRS is computed in a single forward pass without reward labels or rollouts, and it shows that selecting problems with the lowest reasoning‑circuit engagement can outperform random selection on several medium‑difficulty benchmarks. However, the benefit depends on domain, model scale, and reward conditions, indicating that data selection in this setting is regime‑dependent rather than a fixed ranking of problem quality.
The study investigates whether language models can explicitly report constraints they have learned through post‑training fine‑tuning. Using constrained recipe generation with five banned ingredients, the authors compare supervised fine‑tuning (SFT) and Group Relative Policy Optimization (GRPO) against an untrained baseline on a Constraint Awareness Benchmark. Both fine‑tuning methods increase behavioral compliance from 4% to about 90% but reduce explicit constraint reporting and erode retained third‑person knowledge, with GRPO showing more destructive effects. The results suggest that reward‑based signals may suppress constraints context‑independently, and that models fail to enumerate constraints on request even when they can avoid them internally.