arXiv Machine Learning By Chongyu Fan, Pengfei Liu, Jingjia Huang, Sijia Liu, Yi Lin

UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma

Read the original on arXiv Machine Learning →

arXiv:2607. 06987v1 Announce Type: new Abstract: Reinforcement learning (RL) has become the standard paradigm for enhancing the complex reasoning capabilities of large language models (LLMs).

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 10

Stable-MM-R1: Anchoring Multimodal Reasoning Dynamics via Entropy-Guided Stratification

The paper introduces Stable-MM-R1, a framework that stabilizes reinforcement learning for multimodal reasoning by addressing training instability and entropy collapse. It proposes Potential‑Aware Query Mining (PAQM) to filter data toward high‑potential samples and Hybrid Stratified Replay (HSR) to restructure batches using path entropy and reward stratification, reusing stability anchors and hard negatives. The method demonstrates superior performance on complex reasoning tasks compared to strong baselines.

By Yimeng Ye, Shuang Chen, Wenxuan Huang, Manyuan Zhang, Kaituo Feng, Zhangquan Chen, Jiayu Chen, Yucheng Zhou, Yicheng Xiao, Zhiyuan Feng, Tianyu Shi