arXiv Computer Vision By Yunheng Li, Guohong Mu, Hao Li, Shengsheng Qian, Dingwen Zhang, Qibin Hou, Ming-Ming Cheng

Annotations as Rollouts: Efficient and Scalable Reinforcement Learning for Video MLLMs

Read the original on arXiv Computer Vision →

The paper introduces OraRL, a reinforcement learning framework that leverages annotations as oracle rollouts to improve sample efficiency and scalability for video multimodal large language models (MLLMs). By decoupling advantage estimation and employing sign‑balanced pruning, OraRL achieves faster training and better performance across multiple video‑perception benchmarks compared to existing methods. The approach scales from 0.8B to 9B parameters and handles up to 100k prompts, delivering significant gains in temporal mIoU, tracking accuracy, segmentation, and spatial‑intelligence metrics.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv AI
6d ago

Stepwise Intrinsic Rewards for Reasoning in Large Language Models

The paper introduces Stepwise Marginal Information Gain (MIG), an intrinsic process reward that evaluates how each reasoning step of a large language model (LLM) or vision-language model (VLM) improves the likelihood of the reference answer. MIG rewards only new likelihood maxima, preventing duplicate credit, and is combined with outcome, format, and self‑distillation objectives to guide training. Experiments on eight benchmarks show that this method outperforms outcome‑only reinforcement learning and improves accuracy by up to 4.8 points over binary‑reward training, including a 12.6‑point gain on MathVerse and a 12.9‑point advantage on vision‑language transfer at 7B parameters.

By Xiangwei Wang, Wei Wang, Ken Chen, Nanduni Nimalsiri, Sachith Seneviratne, Saman Halgamuge
arXiv AI
Jun 24

video-SALMONN-R$^3$: Learning to ReWatch, ReAsk, and ReAnswer for Efficient Video Understanding

arXiv:2606. 24477v1 Announce Type: cross Abstract: Video large language models (LLMs) are often constrained by computation and memory budgets, leading them to use reduced frame rates and spatial resolutions, which may cause them to miss critical information for question answering (QA).

By Yixuan Li, Guangzhi Sun, Yudong Yang, Wei Li, Zejun MA, Chao Zhang
arXiv Computer Vision
Sep 23

Video-HopChain: Multi-Hop Questions and Confidence-Gated Exploration for Video Reasoning Models

Video-HopChain introduces a new dataset of 22,550 multi‑hop video questions over 13,378 videos, each question consisting of three to six yes/no sub‑questions whose integer answers sum to a verifiable reward. Training a Qwen3‑VL‑8B model with GRPO on this dataset improves performance across eight video‑understanding benchmarks from 55.4 to 57.9, and the addition of Confidence‑Gated Exploration (CGE) raises the mean to 59.3. The authors release the dataset, checkpoint, and training code for further research.

By Trung Nguyen Quang, Yuhao Dong, Shuo Sun, Shuai Liu, Shulin Tian, Kim-Hui Yap, Ziwei Liu