ProcessThinker: Enhancing Multi-modal Large Language Models Reasoning via Rollout-based Process Reward
arXiv:2606. 11209v1 Announce Type: cross Abstract: Visual question answering increasingly requires multi-step reasoning.
Video-HopChain introduces a new dataset of 22,550 multi‑hop video questions over 13,378 videos, each question consisting of three to six yes/no sub‑questions whose integer answers sum to a verifiable reward. Training a Qwen3‑VL‑8B model with GRPO on this dataset improves performance across eight video‑understanding benchmarks from 55.4 to 57.9, and the addition of Confidence‑Gated Exploration (CGE) raises the mean to 59.3. The authors release the dataset, checkpoint, and training code for further research.
arXiv:2606. 11209v1 Announce Type: cross Abstract: Visual question answering increasingly requires multi-step reasoning.
arXiv:2607. 02959v1 Announce Type: cross Abstract: We introduce VSeek, an agentic framework that transforms long-video question answering (LVQA) from a passive, single-pass perception task into a multi-turn retrieval process.
arXiv:2610.01973v1 Announce Type: new Abstract: Reinforcement learning (RL) for video generation usually assigns one scalar reward to an entire sampled video. Yet a video is not uniformly flawed: som...
The paper introduces OraRL, a reinforcement learning framework that leverages annotations as oracle rollouts to improve sample efficiency and scalability for video multimodal large language models (MLLMs). By decoupling advantage estimation and employing sign‑balanced pruning, OraRL achieves faster training and better performance across multiple video‑perception benchmarks compared to existing methods. The approach scales from 0.8B to 9B parameters and handles up to 100k prompts, delivering significant gains in temporal mIoU, tracking accuracy, segmentation, and spatial‑intelligence metrics.
arXiv:2606. 24477v1 Announce Type: cross Abstract: Video large language models (LLMs) are often constrained by computation and memory budgets, leading them to use reduced frame rates and spatial resolutions, which may cause them to miss critical information for question answering (QA).
arXiv:2602. 13602v2 Announce Type: replace-cross Abstract: We present \revise (\underline{Re}asoning with \underline{Vi}deo \underline{S}parsity), a multi-round agent for video question answering (VQA).
arXiv:2510. 17045v2 Announce Type: replace-cross Abstract: Video reasoning using Large Multimodal Models (LMMs) relies on costly reinforcement learning (RL) and verbose chain-of-thought, resulting in substantial computational overhead during both training and inference.
arXiv:2606. 01599v1 Announce Type: new Abstract: Reinforcement learning (RL) for visual reasoning needs scalable, verifiable, and controllable training signals.
The paper introduces VWG-Bench, a benchmark covering nine reasoning dimensions and 38 tasks to evaluate video generative models on symbolic reasoning, physical laws, and goal pursuit. It also presents Vid-PRE, a prompt-rewriting framework that offloads reasoning to a VLM, improving logical performance without changing the generator architecture. Experiments show that current models excel at visual quality but struggle with logic-heavy tasks, while Vid-PRE significantly boosts reasoning across multiple generators.
arXiv:2508. 07683v2 Announce Type: replace-cross Abstract: Video Temporal Grounding (VTG) aims to localize specific video segments corresponding to natural language queries.
The paper introduces Stepwise Marginal Information Gain (MIG), an intrinsic process reward that evaluates how each reasoning step of a large language model (LLM) or vision-language model (VLM) improves the likelihood of the reference answer. MIG rewards only new likelihood maxima, preventing duplicate credit, and is combined with outcome, format, and self‑distillation objectives to guide training. Experiments on eight benchmarks show that this method outperforms outcome‑only reinforcement learning and improves accuracy by up to 4.8 points over binary‑reward training, including a 12.6‑point gain on MathVerse and a 12.9‑point advantage on vision‑language transfer at 7B parameters.
arXiv:2609.36826v1 Announce Type: new Abstract: Video generation models have demonstrated emerging zero-shot capabilities for visual reasoning, perception, and other vision tasks. However, diffusion-...