The paper introduces Stepwise Marginal Information Gain (MIG), an intrinsic process reward that evaluates how each reasoning step of a large language model (LLM) or vision-language model (VLM) improves the likelihood of the reference answer. MIG rewards only new likelihood maxima, preventing duplicate credit, and is combined with outcome, format, and self‑distillation objectives to guide training. Experiments on eight benchmarks show that this method outperforms outcome‑only reinforcement learning and improves accuracy by up to 4.8 points over binary‑reward training, including a 12.6‑point gain on MathVerse and a 12.9‑point advantage on vision‑language transfer at 7B parameters.
By Xiangwei Wang, Wei Wang, Ken Chen, Nanduni Nimalsiri, Sachith Seneviratne, Saman Halgamuge
arXiv:2609.36893v1 Announce Type: new
Abstract: Detailed image captioning requires accurate and comprehensive descriptions of fine-grained visual content, yet caption quality spans factual accuracy,...
By Zhenwen Ji, Lei Jin, Shanyong Wang, Jiaming Lu, Chengqiang Lu, Yi Wu, Yao Hu, Lizhen Cui, Yanyu Xu
arXiv:2609.00591v1 Announce Type: new
Abstract: An image may be worth a thousand words, but most captioning models describe it in only a few. Modern vision-language models produce fluent high-level c...
By Suryaansh Jain, Rahasya Barkur, Vishal G, Ryan Rossi, Franck Dernoncourt, Jack Wang, Koustava Goswami, Nedim Lipka, Puneet Mathur, Samyadeep Basu, Seunghyun Yoon
The paper presents CycleGRPO, a reinforcement learning framework that unifies region understanding and localization for multimodal large language models (MLLMs). By treating the MLLM as both actor and critic, the method generates region captions and immediately grounds them back into spatial coordinates, using a token‑level cycle‑consistency reward that obviates the need for textual ground truths. Experiments on SAMTok demonstrate that CycleGRPO can bootstrap region captioning, VQA, grounded dialogue, and referring segmentation simultaneously, achieving consistent performance gains without task‑specific fine‑tuning.
By Xin Zhang, Haochen Wang, Yikang Zhou, Zhuochen Wang, Xiangtai Li, Robby T. Tan
arXiv:2609.37119v1 Announce Type: cross
Abstract: Recent approaches to reinforcement learning (RL) post-training for large language models increasingly remove the critic to reduce training instabilit...
By Hongyang Li, Xiao Li, Caesar Wu, Said Mammar, Gr\'egoire Danoy, Pascal Bouvry
The paper introduces POISE, a reinforcement learning algorithm that uses a model’s internal states as a value estimator to reduce variance in reinforcement learning with verifiable rewards (RLVR). By employing a lightweight probe that reads internal signals during the forward pass, POISE predicts baselines online and uses a cross‑rollout construction to keep gradients unbiased. Experiments on Qwen3‑4B and OLMo3‑7B‑Instruct‑DPO across six domains show POISE outperforms existing RLVR baselines, offering more stable training and a value model that generalizes across tasks and scales with the policy.
By Yunho Choi, Jongwon Lim, Woojin Ahn, Minjae Oh, Jeonghoon Shim, Yohan Jo
PoEM predicts reinforcement learning outcomes for a new reward function using models already trained on other rewards. If the new reward is a linear combination of existing ones, the new policy’s log-space representation can be expressed as a linear combination of existing log-policies. Even when rewards are not linearly related, log-policies often span a low‑rank subspace, allowing the weighting coefficients to be estimated from reward or basis policy outputs, enabling policy approximation without additional RL training.
By Kimia Hamidieh, Giannis Daras, Antonio Torralba
GrammarRL introduces a label‑free reinforcement learning approach that adapts language models to grammar constraints without annotated data. It optimizes two self‑supervised rewards—direct and reverse—using a Reinforce Leave‑One‑Out objective over grammar‑constrained rollouts, and regularizes toward a frozen base model. Experiments on sign‑language gloss translation, hierarchical text classification, and named entity recognition with Llama models show consistent gains over constrained greedy decoding and competitive performance to beam search while keeping inference cost low.
By Gabriele Tuccio, Antonino Furnari, Aldo Gangemi, Misael Mongiov\`{\i}
arXiv:2610.02015v1 Announce Type: cross
Abstract: Recent advances in LLM reasoning models---driven primarily by the paradigm of post-training via reinforcement learning with verifiable reward (RLVR)-...
By Michael Sullivan, Alexander Koller
arXiv:2606. 11209v1 Announce Type: cross Abstract: Visual question answering increasingly requires multi-step reasoning.
By Jingpei Wu, Xiao Han, Weixiang Shen, Boer Zhang, Zifeng Ding, Volker Tresp
Multi-modal Large Language Models (MLLMs) have achieved remarkable progress in video temporal grounding with reinforcement learning for generating reasoning paths. However, existing models often produce superficial reasoning, which offers limited guidance for precise temporal localization.
Falcon Perception-HD applies reinforcement learning (GRPO) to autoregressive perception models, aligning them directly with precision and recall metrics rather than relying on maximum‑likelihood fine‑tuning. The RL framework introduces reward design for set‑structured outputs and multi‑head sampling control, enabling state‑of‑the‑art performance in very dense scenes (up to 500 objects) and eliminating common issues such as mask repetitions, NMS, and coordinate deduplication. Hybrid self‑annotation pipelines tailored for difficult referring expressions and dense scenes further boost RL training, with improvements observed across all difficulty levels on PBench and SACO‑Gold, and the model preserves object existence knowledge without negative samples.