Rank-Then-Act: Reward-Free Control from Frame-Order Progress
arXiv:2607. 01897v1 Announce Type: cross Abstract: We introduce Rank-Then-Act (RTA), a framework for learning control policies from expert video demonstrations without environment rewards.
arXiv:2607. 01897v1 Announce Type: cross Abstract: We introduce Rank-Then-Act (RTA), a framework for learning control policies from expert video demonstrations without environment rewards.
arXiv:2604. 16557v2 Announce Type: replace Abstract: Current post-training methodologies for adapting Large Vision-Language Models (LVLMs) generally fall into two paradigms: Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL).
The paper introduces OraRL, a reinforcement learning framework that leverages annotations as oracle rollouts to improve sample efficiency and scalability for video multimodal large language models (MLLMs). By decoupling advantage estimation and employing sign‑balanced pruning, OraRL achieves faster training and better performance across multiple video‑perception benchmarks compared to existing methods. The approach scales from 0.8B to 9B parameters and handles up to 100k prompts, delivering significant gains in temporal mIoU, tracking accuracy, segmentation, and spatial‑intelligence metrics.
arXiv:2602. 19313v2 Announce Type: replace-cross Abstract: General-purpose robot learning requires dense, instruction-conditioned feedback that can distinguish meaningful task progress from stalled, failed, or partially completed behavior.
Video-HopChain introduces a new dataset of 22,550 multi‑hop video questions over 13,378 videos, each question consisting of three to six yes/no sub‑questions whose integer answers sum to a verifiable reward. Training a Qwen3‑VL‑8B model with GRPO on this dataset improves performance across eight video‑understanding benchmarks from 55.4 to 57.9, and the addition of Confidence‑Gated Exploration (CGE) raises the mean to 59.3. The authors release the dataset, checkpoint, and training code for further research.
The paper introduces Stepwise Marginal Information Gain (MIG), an intrinsic process reward that evaluates how each reasoning step of a large language model (LLM) or vision-language model (VLM) improves the likelihood of the reference answer. MIG rewards only new likelihood maxima, preventing duplicate credit, and is combined with outcome, format, and self‑distillation objectives to guide training. Experiments on eight benchmarks show that this method outperforms outcome‑only reinforcement learning and improves accuracy by up to 4.8 points over binary‑reward training, including a 12.6‑point gain on MathVerse and a 12.9‑point advantage on vision‑language transfer at 7B parameters.
arXiv:2609.22947v1 Announce Type: new Abstract: Reinforcement learning (RL) is vital for optimizing video generation models, with a robust reward model (RM) serving as the cornerstone. However, exist...
arXiv:2607.08837v4 Announce Type: replace-cross Abstract: Exploration is essential to RL since a policy cannot improve by repeatedly sampling the behaviors it already prefers. Standard methods inject...
Reinforcement learning (RL) is vital for optimizing video generation models, with a robust reward model (RM) serving as the cornerstone. However, existing video reward models often produce unstable sc...
The paper introduces Objective-aware Trajectory Credit Assignment (OTCA), a framework that refines reinforcement learning for diffusion-based visual generation. OTCA decomposes credit across denoising steps and allocates multiple reward signals adaptively, addressing the coarse, uniform reward assignment of existing GRPO pipelines. Experiments demonstrate that OTCA consistently enhances image and video generation quality across various metrics.
arXiv:2606. 29892v1 Announce Type: cross Abstract: Reinforcement learning (RL) has become indispensable for pushing Vision-Language-Action Models (VLAs) beyond static imitation learning.
arXiv:2606. 03937v1 Announce Type: new Abstract: While token-level entropy is commonly recognized as effective for credit assignment in text-only reinforcement learning with verifiable rewards (RLVR), it remains unclear whether this mechanism still holds in visual reasoning.