arXiv AI By Dmitriy Poyarkov, Aleksei Staroverov, Aleksandr I. Panov

Leveraging Offline Supervision for Efficient and Generalizable Reinforcement Learning in Large-Scale Vision-Language-Action Models

Read the original on arXiv AI →

arXiv:2607. 19399v1 Announce Type: cross Abstract: It is commonly observed that online reinforcement learning (RL) produces better-performing strategies than offline methods across a broad range of performance measures.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 7

VLA-Precision: Asymmetric Co-Bootstrapping for Efficient Real-World Online RL of Vision-Language-Action Models

VLA-Precision introduces an efficient real‑world online reinforcement learning framework for vision‑language‑action (VLA) models, featuring the Asymmetric Co‑Bootstrapping (ACoB) algorithm and the ACoB‑Stream architecture. ACoB uses asymmetric co‑bootstrapping across timescales to rapidly improve policy performance while refining value estimates, thereby reducing policy drift. ACoB‑Stream enables large VLA models to run with up to 10.9× higher throughput and computational efficiency, achieving a 98.3 % mean success rate on nine high‑precision chemistry tasks in under 46 minutes per task.

By Chenyu Su, Zhaolong Shen, Yuan Qian, Chen Qian, Rui Zhang, Feng Yan, Weixing Chen, Fei Zhang, Jiamin Wang, Shuang Cong, Weiwei Shang