arXiv Machine Learning

Reward Granularity in RLVR: Comparing Process and Outcome Reward Structures for Mathematical Reasoning in Small Language Models

arXiv:2607. 02869v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a promising paradigm for improving mathematical reasoning in language models.

arXiv Machine Learning
Sep 4

Gradients Know What Outcomes Don't: Unlocking Reinforcement Learning for LLM Reasoning with Gradient-Aligned Rewards

The paper introduces Gradient-Aligned Reward (GAR), a reinforcement learning technique that uses truncated backpropagation to generate a compact gradient vector for each rollout and compares it to an expert-anchor gradient via cosine similarity. This dense, reasoning-aware reward improves large language model chain-of-thought reasoning on math benchmarks and transfers to other tasks without domain‑specific data, while adding less than 9% computational overhead. GAR outperforms existing baselines such as GRPO on Qwen3-4B and Qwen3-8B models.

By Leqi Zheng, Jinbo Su, Fang Niu, Chaokun Wang, Weiping Wang, Jiajun Zhang, Shannan Yan, Jie Wu, Zhaolu Kang, Rong Fu, Hang Zhang
arXiv AI
Sep 25

Reinforcement Learning with Verifiable Rewards for Small Search Agents

The paper introduces Reinforcement Learning with Verifiable Rewards (RLVR) applied to small search agents, specifically training a Qwen3.5-0.8B model with Group Relative Policy Optimization and an interleaved Wikipedia-search tool on the MuSiQue dataset. Experiments varying reward shapes across three seeds show that RLVR can achieve a 3.8‑fold improvement over an untrained baseline, with the best run reaching a 0.352 average exact match. The study finds that the sparse exact‑match reward, standard in larger models, performs poorly for small models, indicating that reward design must be tailored rather than scaled down from large‑model recipes.

By Gaurisankar Jayadas, Aske Plaat, \'Alvaro Serra-G\'omez, Sandheep P
arXiv AI
Sep 15

HISPO: Hierarchical Importance-Sampling Policy Optimization with Entropy-Derived Segments

HISPO introduces a segment‑level policy‑optimization technique for reinforcement learning with verifiable rewards, creating entropy‑derived contiguous segments during rollout and applying clipped importance‑sampling correction at this granularity. It bridges the gap between token‑level and sequence‑level corrections, offering a middle‑ground approach for credit assignment in long‑form mathematical reasoning. Evaluated on Qwen3‑1.7B‑Base across six benchmarks, HISPO consistently outperforms or matches the strongest baselines in Pass@8 and Acc@8 metrics, notably improving AIME25 scores over GRPO and GSPO.

By Quoc-Vinh Lai-Dang, Hyo-Sang Shin