arXiv Machine Learning

TARC: Time-Adaptive Robotic Control

TARC (Time‑Adaptive Robotic Control) is a reinforcement‑learning framework that lets a policy predict both a control action and how long it should be applied, thereby learning temporally extended actions. By optimizing task performance under constraints on the number of control switches, TARC can adapt its control rate online, using high‑frequency feedback only when necessary. Experiments on a high‑speed RC car, a Unitree Go1 quadruped, and a vision‑language action model show that TARC matches the performance of high‑frequency discrete‑time controllers while operating at less than half their control frequency.

arXiv AI
Aug 10

Learning to Walk With Less: A Dyna-Style Approach to Quadrupedal Locomotion

arXiv:2509. 06296v2 Announce Type: replace-cross Abstract: Traditional on-policy reinforcement learning (RL) controllers for quadrupedal locomotion often suffer from low data efficiency, requiring millions of interactions with simulated environments to achieve stable control.

By Francisco Affonso, Felipe Tommaselli, Jo\~ao H. Al\'essio, Vivian S. Medeiros, Mateus V. Gasparino, Girish Chowdhary, Marcelo Becker
arXiv Machine Learning
Sep 17

Reinforcement Learning for Real-Time Vision-Language-Action Policies

The paper presents Real‑Time EXPO‑FT, a reinforcement learning framework that fine‑tunes large Vision‑Language‑Action models for real‑time robotic control. It separates slow, expressive action generation from fast, reactive edits, allowing a lightweight policy to adjust actions based on the latest observation. Experiments on the Kinetix benchmark and four dynamic real‑world tasks show that Real‑Time EXPO‑FT achieves superior performance, improving policy success rates from 42% to 97% with only ten minutes of online data and no human intervention.

By Perry Dong, Kuo-Han Hung, Dorsa Sadigh, Chelsea Finn
arXiv AI
Aug 19

Q-Learning With World Models

The paper introduces QWM, a framework that integrates world models with standard Q‑learning to perform test‑time search over imagined trajectories. By training the policy and value function solely on real transitions, QWM avoids compounding model bias while still benefiting from predictive search. Experiments on the Robomimic and LIBERO manipulation benchmarks show that QWM outperforms strong prior state‑of‑the‑art methods in both sample efficiency and performance.

By Perry Dong, Yueru Jia, Chelsea Finn, Dorsa Sadigh
arXiv AI
Sep 24

BEE: Intervention-Adaptive Real-World Reinforcement Learning with Vision-Language-Action Models

arXiv:2609.27450v1 Announce Type: cross Abstract: Vision-language-action (VLA) models handle long-horizon manipulation, yet success hinges on a few precision-critical phases where millimeter-scale er...

By Weihui Zhao, Xiaohan Yan, Zunian Wan, Xuan Du, Zhaozhan Chi, Jianbo Mao, Ruipu Wu, Rushuai Yang, Houlin Li, Shukai Yang, Jing Wu, Yuxiang Yan, Yongcheng Liu, Chuankang Li, Guanghui Ren, Wei Shan, Maoqing Yao