arXiv AI

AtomBridge: Agentic VLA Inference Plugin for Long-Horizon Tasks in Scientific Experiments

arXiv:2602. 09430v2 Announce Type: replace-cross Abstract: Robotic laboratories play a critical role in autonomous scientific discovery by enabling scalable, continuous experimental execution.

arXiv AI
Sep 10

AtomicVLA: Unlocking the Potential of Atomic Skill Learning in Robots

AtomicVLA is a unified planning-and-execution framework that generates task-level plans, atomic skill abstractions, and fine-grained actions for robotic manipulation. It builds a scalable atomic skill library using a Skill‑Guided Mixture‑of‑Experts (SG‑MoE) and a flexible routing encoder that assigns new skills to dedicated experts, enabling continual learning. Experiments show that AtomicVLA outperforms baseline models on both simulated and real‑world long‑horizon tasks, achieving significant improvements in task performance and learning efficiency.

By Likui Zhang, Tao Tang, Zhihao Zhan, Xiuwei Chen, Zisheng Chen, Jianhua Han, Jiangtong Zhu, Pei Xu, Hang Xu, Hefeng Wu, Liang Lin, Xiaodan Liang
arXiv AI
Jun 12

LabVLA: Grounding Vision-Language-Action Models in Scientific Laboratories

arXiv:2606. 13578v1 Announce Type: cross Abstract: Scientific laboratories increasingly rely on AI systems to reason about experiments, but the physical act of doing science remains largely outside their reach.

By Baochang Ren, Xinjie Liu, Xi Chen, Yanshuo Liu, Chenxi Li, Daqi Gao, Zeqin Su, Jintao Xing, Zirui Xue, Rui Li, Xiangyu Zhao, Shuofei Qiao, Minting Pan, Wangmeng Zuo, Lei Bai, Dongzhan Zhou, Ningyu Zhang, Huajun Chen
arXiv AI
Aug 19

EXPO-FT: Sample-Efficient Reinforcement Learning Finetuning for Vision-Language-Action Models

EXPO-FT is a system that enables stable, sample‑efficient reinforcement learning fine‑tuning of pretrained Vision‑Language‑Action (VLA) policies. It achieves perfect success on a range of manipulation tasks—such as routing string lights, striking a pool ball, and inserting a flower into a wine bottle—using only about 19.1 minutes of online robot data. The approach outperforms both RL-from-scratch and existing VLA fine‑tuning methods, and the authors provide an open‑source codebase to support wider adoption.

By Perry Dong, Kuo-Han Hung, Tian Gao, Dorsa Sadigh, Chelsea Finn
arXiv Machine Learning
Aug 26

Learning to Act While Waiting: RL Finetuning of Generalist Robot Policies Under Inference Latency

The paper introduces ARLI, a latency‑aware framework that enables reinforcement learning fine‑tuning of large generalist robot policies despite inference delays. ARLI combines asynchronous inference with state augmentations—incorporating committed actions and mid‑inference observations—to restore near‑Markovian dynamics and maintain reactivity. Experiments on simulated and real‑world manipulation tasks show that ARLI allows effective policy improvement under latency, outperforming standard RL even in no‑latency scenarios.

By Brian Zhu (Siemens), Momen Khalil (Siemens), E Harrison (UC Berkeley), Emanuele Poggi (Siemens), Philipp Schmitt (Siemens), Bernd Kast (Siemens), Philine Meister (Siemens), Pranav Atreya (UC Berkeley), Qiyang Li (UC Berkeley), Finn Ferchau (Siemens), Cesar Colmenero (Siemens), Yash Shahapurkar (Siemens), Gokul Narayanan (Siemens), Melih Erdogan (Siemens), Kai Wurm (Siemens), Georg von Wichert (Siemens), Oier Mees (Microsoft, ETH Zurich, UC Berkeley), Eugen Solowjow (Siemens), Andrew Wagenmaker (UC Berkeley), Sergey Levine (UC Berkeley)
arXiv AI
Jul 28

A Few Words Go a Long Way: Language Guided Robot Policy Synthesis

arXiv:2607. 23784v1 Announce Type: cross Abstract: While vision-language-action models have demonstrated impressive zero-shot manipulation capabilities, they remain fundamentally black box policies that are difficult to interpret, adapt, or correct when they inevitably fail.

By Daphne Chen, Archit Ritesh Jain, Eric Goossen, Emma Romig, Michael Murray, Nick Walker, Maya Cakmak