arXiv AI By Wooil Jung

Dropout-GRPO: Variational Stochasticity for Continuous Latent Reasoning

Read the original on arXiv AI →

arXiv:2606. 10184v1 Announce Type: cross Abstract: Group Relative Policy Optimization (GRPO) relies on the diversity of $K$ rollouts within each group; otherwise, the group-mean advantage $A^{(k)} = r^{(k)} - \mu_r$ collapses to zero.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 14

Sampling via Decision-Flow: Training-Free Extraction of Improved Latent Reasoning Paths in Large Language Models

The paper introduces Decision-Flow Sampling (DF‑Sample), a training‑free, data‑free inference framework that builds a hierarchical reasoning tree, evaluates entire trajectories, and back‑propagates utilities to guide branching decisions. Unlike local step‑wise sampling, DF‑Sample explicitly assesses global paths, enabling it to recover high‑quality, low‑probability reasoning chains that standard decoding misses. On the GPQA benchmark, DF‑Sample attains 45.6% accuracy, outperforming power sampling (38.9%) and GRPO (39.9%) and consistently surpassing baselines across multiple models and benchmarks, demonstrating significant latent reasoning potential in pretrained LLMs.

By Zhendong Mi, Shaoyi Huang
arXiv Machine Learning
Aug 28

Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO

The paper investigates Evolution Strategies (ES) as a memory‑efficient post‑training method for large language model (LLM) reasoning. It demonstrates that ES outperforms Group Relative Policy Optimization (GRPO) by achieving broader reasoning coverage, improving Pass@K metrics, and avoiding entropy collapse. The study also reveals that ES’s performance gains stem from sparse, high‑magnitude parameter updates, do not cause catastrophic forgetting, and can be combined with GRPO in a sequential training strategy.

By Yunpeng Ba, Zhi Zheng, Yue Xie, Jiaqing Li, Xialiang Tong, Tao Zhong, Mingxuan Yuan, Zhichao Lu, Xuyang Wu, Zhenkun Wang
arXiv AI
2d ago

Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States

The paper introduces POISE, a reinforcement learning algorithm that uses a model’s internal states as a value estimator to reduce variance in reinforcement learning with verifiable rewards (RLVR). By employing a lightweight probe that reads internal signals during the forward pass, POISE predicts baselines online and uses a cross‑rollout construction to keep gradients unbiased. Experiments on Qwen3‑4B and OLMo3‑7B‑Instruct‑DPO across six domains show POISE outperforms existing RLVR baselines, offering more stable training and a value model that generalizes across tasks and scales with the policy.

By Yunho Choi, Jongwon Lim, Woojin Ahn, Minjae Oh, Jeonghoon Shim, Yohan Jo