arXiv AI By Yingxuan Zhuang, Miao Pan, Wangjie Gan, Jingxiao Yang, Fan Wang, Weiming Liu, Cheng Tan, Xuhong Zhang, Jintao Chen

DEEPO: Dual-Entropy Enhanced Policy Optimization for Hallucination in MLLMs

Read the original on arXiv AI →

The paper introduces DEEPO, a Dual-Entropy Enhanced Policy Optimization method designed to mitigate hallucination in multimodal large language models (MLLMs). It addresses two weaknesses in reinforcement learning: (1) hard queries with high semantic entropy produce uniformly wrong samples, erasing advantage signals, and (2) confident-but-wrong tokens become invisible to gradients as the policy sharpens. DEEPO combines semantic‑entropy‑triggered expert prefixes to inject grounded continuations and Renyi preconditioning to counter logit saturation, yielding significant hallucination reduction while maintaining accuracy and training stability.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 10

Stable-MM-R1: Anchoring Multimodal Reasoning Dynamics via Entropy-Guided Stratification

The paper introduces Stable-MM-R1, a framework that stabilizes reinforcement learning for multimodal reasoning by addressing training instability and entropy collapse. It proposes Potential‑Aware Query Mining (PAQM) to filter data toward high‑potential samples and Hybrid Stratified Replay (HSR) to restructure batches using path entropy and reward stratification, reusing stability anchors and hard negatives. The method demonstrates superior performance on complex reasoning tasks compared to strong baselines.

By Yimeng Ye, Shuang Chen, Wenxuan Huang, Manyuan Zhang, Kaituo Feng, Zhangquan Chen, Jiayu Chen, Yucheng Zhou, Yicheng Xiao, Zhiyuan Feng, Tianyu Shi