Aligning Thoughts with Answers: Probability Rewards to Tame Thinking Drift
Read the original on Hugging Face Trending Papers →The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The paper introduces Rita, a reinforcement learning framework that addresses thinking drift in vision‑language models by enforcing consistency between reasoning and answers. Rita employs two reasoning‑label‑free rewards—thinking and consistency rewards—derived from the conditional probability of reference answers, and uses a difficulty‑aware data filtering strategy to select informative samples for training. Experiments on EgoIntention and RefEgo‑Int benchmarks demonstrate that Rita outperforms both supervised fine‑tuning and vanilla RL‑fine‑tuned approaches.
arXiv:2610.02015v1 Announce Type: cross Abstract: Recent advances in LLM reasoning models---driven primarily by the paradigm of post-training via reinforcement learning with verifiable reward (RLVR)-...
arXiv:2606. 29984v1 Announce Type: new Abstract: Reinforcement Learning (RL) is an important paradigm for improving the reasoning capabilities of Vision-Language Models (VLMs).
The paper introduces Stepwise Marginal Information Gain (MIG), an intrinsic process reward that evaluates how each reasoning step of a large language model (LLM) or vision-language model (VLM) improves the likelihood of the reference answer. MIG rewards only new likelihood maxima, preventing duplicate credit, and is combined with outcome, format, and self‑distillation objectives to guide training. Experiments on eight benchmarks show that this method outperforms outcome‑only reinforcement learning and improves accuracy by up to 4.8 points over binary‑reward training, including a 12.6‑point gain on MathVerse and a 12.9‑point advantage on vision‑language transfer at 7B parameters.
arXiv:2606. 11209v1 Announce Type: cross Abstract: Visual question answering increasingly requires multi-step reasoning.
arXiv:2608. 03875v1 Announce Type: cross Abstract: Designing effective reward functions remains a major bottleneck in Reinforcement Learning (RL).