ProcessThinker: Enhancing Multi-modal Large Language Models Reasoning via Rollout-based Process Reward
arXiv:2606. 11209v1 Announce Type: cross Abstract: Visual question answering increasingly requires multi-step reasoning.
arXiv:2607. 23700v1 Announce Type: new Abstract: Multimodal large language models exhibit capabilities on reasoning tasks, yet often produce flawed intermediate steps while yielding correct final answers.
arXiv:2606. 11209v1 Announce Type: cross Abstract: Visual question answering increasingly requires multi-step reasoning.
arXiv:2609.21675v1 Announce Type: new Abstract: Despite the remarkable progress in Multimodal Large Language Models (MLLMs), prevailing Chain-of-Thought (CoT) paradigms remain confined to the natural...
arXiv:2608. 08326v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has emerged as an effective approach for improving multimodal reasoning.
arXiv:2606. 15160v1 Announce Type: cross Abstract: Reasoning capabilities of multimodal large language models (MLLMs) have improved considerably in recent years.
arXiv:2610.01892v1 Announce Type: cross Abstract: Multimodal agents commonly generate free-form reasoning before each action. For small models, limited model capacity can result in lengthy reasoning...
arXiv:2607. 19450v1 Announce Type: cross Abstract: Large-scale online reinforcement learning (RL) is the predominant means of eliciting advanced abilities including long-term reasoning and agentic tool use in large language models (LLMs).
Reinforcement learning with verifiable rewards (RLVR) has emerged as an effective approach for improving multimodal reasoning. However, most existing methods evaluate an entire response using a binary reward based only on final-answer correctness, thereby discarding the supervision available in intermediate reasoning steps.
The paper introduces Echo-GRPO, a method that rewrites privileged reasoning traces into a model’s own idiolect to align off‑policy supervision with the student policy’s vocabulary. By preserving semantics through Dual‑Reference Decoding, Echo‑GRPO mitigates gradient clipping on critical reasoning tokens and improves reasoning distillation. The approach is instantiated as VideoEcho‑R1 for video reasoning, yielding consistent gains across multiple multimodal LLM backbones and benchmarks, and it can be applied as a plug‑in to both RL and supervised fine‑tuning frameworks.
arXiv:2606. 29984v1 Announce Type: new Abstract: Reinforcement Learning (RL) is an important paradigm for improving the reasoning capabilities of Vision-Language Models (VLMs).
CRYSTAL is a diagnostic benchmark comprising 6,372 multimodal reasoning instances that assess models through verifiable intermediate steps. It introduces two metrics—Match F1 and Ordered Match F1—to evaluate step-level precision, recall, and order. The benchmark, built via a Delphi-inspired pipeline with four independent MLLMs, reveals systematic failures in current models, such as cherry‑picking and disordered reasoning, and proposes a Causal Process Reward and CPR‑Curriculum to improve reasoning performance.
arXiv:2609.08025v1 Announce Type: new Abstract: Reasoning agents increasingly rely on external tools such as web search to answer complex queries. Reinforcement learning (RL) finetuning algorithms su...
CanvasAnneal is a curriculum‑guided reinforcement learning framework designed to improve Diffusion Language Models (DLMs) on complex reasoning and tool‑use tasks. It starts training by injecting reasoning traces from a stronger teacher model into the diffusion canvas, then gradually reduces this guidance so the model learns to generate reasoning independently. Experiments on mathematical reasoning and tool‑use benchmarks show that CanvasAnneal outperforms standard diffusion RL methods such as diffu‑GRPO on tasks like MATH500, Countdown, and Tau2, and accelerates reward improvement, though the gains vary by task.