arXiv Machine Learning

Rollout Total Correlation for Deep Reinforcement Learning

The paper proposes a method for learning task-relevant representations in deep reinforcement learning by maximizing rollout total correlation, which captures the correlation among all learned representations and actions across entire trajectories. It introduces two complementary lower bounds—one generative and one discriminative—along with chunk‑wise mini‑batching to improve this objective, and also proposes an intrinsic reward derived from the learned representation to enhance exploration. Experiments on challenging image‑based simulated control tasks demonstrate improved sample efficiency and robustness to white noise and natural video backgrounds compared to leading baselines.

arXiv Computer Vision
Sep 7

Persistent Robot World Models: Stabilizing Multi-Step Rollouts via Reinforcement Learning

The paper introduces a reinforcement learning post‑training scheme that trains robot world models on their own autoregressive rollouts, using a contrastive RL objective adapted from diffusion models. It also proposes a training protocol that compares multiple variable‑length futures, a multi‑view visual fidelity reward, and demonstrates state‑of‑the‑art rollout fidelity on the DROID dataset, outperforming baselines on LPIPS, SSIM, and human preference tests.

By Jai Bardhan, Patrik Drozdik, Josef Sivic, Vladimir Petrik
arXiv Machine Learning
6d ago

Learning from Mixed-Quality Deployment Experience for Robot Manipulation

The paper introduces Predictive Action Chunk Learning (PACL), a method for improving robot manipulation policies using mixed-quality deployment experience. PACL first trains a predictive chunk-level critic to evaluate temporally extended action sequences, then uses the critic’s quality estimates to guide a diffusion actor that learns from both successful and failed rollouts. Experiments on simulated and real robots demonstrate that PACL consistently enhances pretrained policies and outperforms strong imitation learning and offline reinforcement learning baselines.

By Yangang Ren, Yujie Yan, Zirui Li, Jiaming Guo, Di Zeng, Ji Tao, Lan Yu, Xuesong Tian, Chen Lv
arXiv Computer Vision
Sep 7

Learning to Credit the Right Steps: Objective-aware Process Optimization for Visual Generation

The paper introduces Objective-aware Trajectory Credit Assignment (OTCA), a framework that refines reinforcement learning for diffusion-based visual generation. OTCA decomposes credit across denoising steps and allocates multiple reward signals adaptively, addressing the coarse, uniform reward assignment of existing GRPO pipelines. Experiments demonstrate that OTCA consistently enhances image and video generation quality across various metrics.

By Rui Li, Ke Hao, Yuanzhi Liang, Haibin Huang, Chi Zhang, Yun Gu, Xuelong Li
arXiv Computer Vision
Sep 10

VideoTIR: Accurate Understanding for Long Videos with Efficient Tool-Integrated Reasoning

VideoTIR introduces a reinforcement‑learning approach to improve long‑video understanding by encouraging multimodal large language models to use comprehensive multi‑level toolkits efficiently. It combines Zero‑RL and SFT cold‑starting strategies to help models retrieve and focus on meaningful video segments, images, and regions, thereby reducing hallucinations. The method includes Toolkit Action Grouped Policy Optimization (TAGPO) to streamline tool‑calling and a sandbox‑based trajectory synthesis framework for high‑quality data, achieving strong results on three long‑video QA benchmarks.

By Zhe Gao, Shiyu Shen, Taifeng Chai, Weinong Wang, Haotian Xu, Xing Wu, Wenbin Li, Qi Fan, Yang Gao, Dacheng Tao
arXiv Computer Vision
1d ago

RoXDrive: Closed-Loop Reinforcement Learning for End-to-End Autonomous Driving via Action-Faithful Rollouts

arXiv:2609.36851v1 Announce Type: new Abstract: End-to-end autonomous driving policies are commonly trained via imitation learning on logged demonstrations without observing the consequences of their...

By Hongbin Lin, Chaoda Zheng, Yiming Yang, Xiangyu Li, Shijia Chen, Jinhao Deng, Kangjie Chen, Dongbin Zhang, Jie Feng, Yu Zhang, Xianming Liu, Shuguang Cui, Boyang Wang, Zhen Li