arXiv Machine Learning By Naoki Nishikawa, Taiji Suzuki

Reinforcement Learning for Hierarchical Reasoning Rewards: Minimax-Optimal Rates with Transformers

Read the original on arXiv Machine Learning →

The paper studies reinforcement learning (RL) for post‑training language models on reasoning tasks, focusing on a hierarchical reward structure where each component becomes relevant only after the previous ones are resolved. It demonstrates that a Transformer‑based actor–critic algorithm that alternates between KL‑regularized policy sampling, critic fitting, and policy updates achieves minimax‑optimal rates in query budget and regularization strength, and is optimal for a fixed number of prompts. In contrast, sampling from a fixed reference distribution, as used in offline reward modeling, only yields a logarithmic regret decay, highlighting the advantage of on‑policy exploration in concentrating on high‑reward regions.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
6d ago

Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States

The paper introduces POISE, a reinforcement learning algorithm that uses a model’s internal states as a value estimator to reduce variance in reinforcement learning with verifiable rewards (RLVR). By employing a lightweight probe that reads internal signals during the forward pass, POISE predicts baselines online and uses a cross‑rollout construction to keep gradients unbiased. Experiments on Qwen3‑4B and OLMo3‑7B‑Instruct‑DPO across six domains show POISE outperforms existing RLVR baselines, offering more stable training and a value model that generalizes across tasks and scales with the policy.

By Yunho Choi, Jongwon Lim, Woojin Ahn, Minjae Oh, Jeonghoon Shim, Yohan Jo
Hugging Face Trending Papers
Jun 24

MiniOpt: Reasoning to Model and Solve General Optimization Problems with Limited Resources

Achieving strong optimization generalization across diverse optimization problems while requiring limited training resources remains a challenging problem for optimization-oriented large language models (LLMs). Existing approaches typically rely on large-scale supervised datasets, costly reasoning annotations, and expensive intermediate step verification, resulting in substantial training overhead.