arXiv Machine Learning

UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma

arXiv:2607. 06987v1 Announce Type: new Abstract: Reinforcement learning (RL) has become the standard paradigm for enhancing the complex reasoning capabilities of large language models (LLMs).

arXiv Machine Learning
Sep 10

Stable-MM-R1: Anchoring Multimodal Reasoning Dynamics via Entropy-Guided Stratification

The paper introduces Stable-MM-R1, a framework that stabilizes reinforcement learning for multimodal reasoning by addressing training instability and entropy collapse. It proposes Potential‑Aware Query Mining (PAQM) to filter data toward high‑potential samples and Hybrid Stratified Replay (HSR) to restructure batches using path entropy and reward stratification, reusing stability anchors and hard negatives. The method demonstrates superior performance on complex reasoning tasks compared to strong baselines.

By Yimeng Ye, Shuang Chen, Wenxuan Huang, Manyuan Zhang, Kaituo Feng, Zhangquan Chen, Jiayu Chen, Yucheng Zhou, Yicheng Xiao, Zhiyuan Feng, Tianyu Shi
Hugging Face Trending Papers
Jul 30

LEEPS: Latent-Guided Explore-Exploit Prompt Sampling for Efficient RLVR in Large Language Models

Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models, but prompt groups with identical rollout rewards consume generation budget without effective learning signals. Pre-rollout prompt selection can reduce this waste by screening prompts before rollout generation.