arXiv Computer Vision By Yudong Han, Yong Wang, Zaiquan Yang, Liang Lin, Chongyang Tao, Xiangxiang Chu, Liyuan Pan

Beyond Entropy: Self-Diagnostic Multi-Role Token Optimization for Video Reasoning

Read the original on arXiv Computer Vision →

The paper introduces DyCPO, a co‑evolutionary framework that jointly optimizes token selection and adaptive counterfactual intervention for video reasoning. It builds a multi‑role dependence metric to balance visual exploration with answer‑relevance mining, suppressing filler tokens and spurious visual noise. By deriving counterfactual signals from the model’s own successful and failed rollouts, DyCPO enables self‑diagnostic analysis and co‑evolution of the optimization objective with the policy, leading to consistent performance gains on complex video reasoning benchmarks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv AI
Jun 3

Entropy Is Not Enough: Unlocking Effective Reinforcement Learning for Visual Reasoning via Vision-Anchored Token Selection

arXiv:2606. 03937v1 Announce Type: new Abstract: While token-level entropy is commonly recognized as effective for credit assignment in text-only reinforcement learning with verifiable rewards (RLVR), it remains unclear whether this mechanism still holds in visual reasoning.

By Senjie Jin, Peixin Wang, Boyang Liu, Xiaoran Fan, Shuo Li, Zhiheng Xi, Jiazheng Zhang, Yuhao Zhou, Tao Gui, Qi Zhang, Xuanjing Huang
arXiv AI
Sep 17

Anchoring What Matters: A Dual-Level Learning Framework for Visually-Grounded Multimodal Reasoning

The paper introduces PIVOT, a dual-level learning framework designed to improve visually-grounded multimodal reasoning in large vision-language models. PIVOT employs a self‑calibrated experience replay mechanism to selectively reuse valuable visual reasoning trajectories, and a vision‑guided advantage allocation scheme that assigns extra rewards to tokens with strong visual support. Experiments on multiple benchmarks show that PIVOT enhances the multimodal reasoning performance of these models.

By Xinxin Song, Siyuan Li, Tingxiong Xiao, Jinli Suo
arXiv Computer Vision
Aug 27

Boosting Reasoning in Large Multimodal Models via Activation Replay

The paper introduces Activation Replay, a training‑free method that improves reasoning in post‑trained large multimodal models (LMMs) by replaying low‑entropy activations from the base model’s input context. It shows that Reinforcement Learning with Verifiable Rewards (RLVR) shifts low‑entropy activations and that modulating these activations enhances reasoning across tasks such as mathematics, visual agents, and video reasoning. Experiments demonstrate that Activation Replay outperforms alternatives like high‑entropy replay or direct cross‑model intervention, boosting Pass@K and broadening RLVR’s reasoning coverage.

By Yun Xing, Xiaobin Hu, Qingdong He, Jiangning Zhang, Shuicheng Yan, Shijian Lu, Yu-Gang Jiang
arXiv AI
Aug 21

VISD: Enhancing Video Reasoning via Structured Self-Distillation

arXiv:2605. 06094v5 Announce Type: replace-cross Abstract: Training VideoLLMs for complex reasoning remains challenging due to sparse sequence level rewards and the lack of fine grained credit assignment over long, temporally grounded reasoning trajectories.

By Hao Lin, Kunyang Lv, Xu Jiang, Jingqi Tian, Zhongjing Du, Jiayu Ding, Qiaoman Zhang, Hongbo Jin
arXiv AI
Aug 11

SCOUT: Self-Checking and Recovery-Aware Tool-Thought Agents for Ultra-Long Egocentric Video Reasoning

arXiv:2608. 07959v1 Announce Type: new Abstract: Ultra-long egocentric video understanding requires reasoning over temporally sparse evidence distributed across hours or days, challenging current multimodal models with limited context and the grounding of key video segments.

By Keyang Zhong, Kuo Wang, Peng Liu, Quanlong Zheng, Junlin Xie, Zhijia Liang, Yanhao Zhang, Guanbin Li