VLA-Precision introduces an efficient real‑world online reinforcement learning framework for vision‑language‑action (VLA) models, featuring the Asymmetric Co‑Bootstrapping (ACoB) algorithm and the ACoB‑Stream architecture. ACoB uses asymmetric co‑bootstrapping across timescales to rapidly improve policy performance while refining value estimates, thereby reducing policy drift. ACoB‑Stream enables large VLA models to run with up to 10.9× higher throughput and computational efficiency, achieving a 98.3 % mean success rate on nine high‑precision chemistry tasks in under 46 minutes per task.
By Chenyu Su, Zhaolong Shen, Yuan Qian, Chen Qian, Rui Zhang, Feng Yan, Weixing Chen, Fei Zhang, Jiamin Wang, Shuang Cong, Weiwei Shang
Online Self-Weighted Fine‑Tuning (OSW‑FT) augments standard supervised fine‑tuning by adding online, trajectory‑level weighting: for each query the model estimates its current success rate from a small number of inference‑only rollouts and rescales the SFT loss accordingly. The method keeps the optimization direction anchored to the expert trajectory while adapting the update magnitude online, and it is shown to be unbiased for any finite rollout count with a convergence analysis. Across Qwen3 models from 0.6B to 4B, OSW‑FT consistently outperforms plain SFT on challenging benchmarks such as AIME, achieving a favorable compute‑performance trade‑off with only two online rollouts.
By Haiquan Wen, Yiwei He, Bei Peng, Guangliang Cheng
arXiv:2605. 29032v2 Announce Type: replace Abstract: Model-based reinforcement learning (MBRL) agents typically learn world models by minimizing predictive loss.
By Christoph Dann, Yishay Mansour, Mehryar Mohri
arXiv:2607. 01612v1 Announce Type: new Abstract: Training large language models (LLMs) with reinforcement learning (RL) has significantly advanced their performance on reasoning and question-answering tasks.
By Xuqing Yang, Yi Yuan, Shanzhe Lei, Xuhong Wang
arXiv:2509. 11259v2 Announce Type: replace-cross Abstract: Recent advancements in machine learning have largely been driven by foundation models (FMs) trained on large, diverse datasets, enabling them to generalize effectively to new, related tasks.
By David Schiff, Ofir Lindenbaum, Yonathan Efroni
The paper introduces Solver-Gradient Guided Reinforcement Learning (SG‑RL), a method that augments standard RL with bounded gradients from a differentiable MPC solver to adapt cost‑function weights online. SG‑RL integrates solver‑gradient guidance into PPO through actor‑update scaling, policy loss, advantage estimation, and value‑function learning, achieving comparable or superior closed‑loop performance while requiring up to 70.6% fewer samples. Experiments on two autonomous racing platforms with intentional model mismatch demonstrate that SG‑RL outperforms both RL and gradient‑based policy learning baselines and generalizes zero‑shot to unseen environments.
By Baha Zarrouki, Arslan Thobani, Jasper Hoffmann, Mattia Piccinini, Rudolf Reiter, Felix Jahncke, S\'ebastien Gros, Davide Scaramuzza, Johannes Betz
The paper introduces DCRL (Divide-and-Conquer RL), a method that recursively decomposes offline goal-conditioned reinforcement learning trajectories into a balanced binary tree. By training values from the leaves up to the root, DCRL avoids noisy max-based backups and reduces bootstrap depth from linear to logarithmic, thereby limiting error accumulation. Experiments on diverse goal-reaching tasks show that DCRL outperforms prior flat offline GCRL methods, achieving a higher average score on the most challenging long-horizon OGBench tasks.
By Hyeonseong Jeon, Youngwoon Lee
arXiv:2606. 18284v1 Announce Type: cross Abstract: The limiting resource for training agents via reinforcement learning (RL) is increasingly frontier task supply: valid, solvable tasks just difficult enough to train the current model.
By Lorenz Wolf, Connor Watts, Roger Creus Castanyer, Geoffrey Bradway, Maxwill Lin, Augustine N. Mavor-Parker, Matthew Daborn-Sargent
arXiv:2606. 02194v1 Announce Type: new Abstract: Distilling expert demonstration data into large generative models using behavioral cloning is a scalable approach to learning capable policies for robotic control, particularly for dexterous manipulation.
By Christian Scherer, Joe Watson, Theo Gruner, Daniel Palenicek, Ingmar Posner, Jan Peters
arXiv:2608.24479v1 Announce Type: new
Abstract: Massively parallel simulation changes the data regime in which off-policy reinforcement learning (RL) is trained, challenging stabilizers designed for...
By Zihao Wu, Hongyao Tang, Yi Ma, Huizhong Song, Pengyi Li, Yifu Yuan, Fei Ni, Jinyi Liu, Wei Wei, Jianrong Wang, Yan Zheng, Jianye Hao
arXiv:2606. 27369v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) for training LLMs typically rely on ground-truth answers to assign rewards, limiting their applicability to tasks where the ground-truth solution is unknown.
By Yingyu Lin, Qiyue Gao, Nikki Lijing Kuang, Xunpeng Huang, Kun Zhou, Tongtong Liang, Zhewei Yao, Yi-An Ma, Yuxiong He
arXiv:2601. 15141v2 Announce Type: replace Abstract: Agentic Reinforcement Learning (RL) has empowered Large Language Models (LLMs) to utilize tools like Python interpreters for complex problem-solving.
By Tianshi Xu, Yuteng Chen, Meng Li