arXiv AI By Yoojin Oh, Jeongsol Kim, Yeonwoo Seo, Jangho Park, Seonghyun Jin, Sunwoo Park, Youngmin Kim, Youngjun Jun, Kyumin Choi, Jong Chul Ye

FastOPD: On-Policy Distillation for Lightweight VLA Deployment

Read the original on arXiv AI →

FastOPD is a framework that distills large Vision‑Language‑Action (VLA) models into lightweight versions by using on‑policy distillation with a flow map and a self‑consistency objective. The method trains a compact student to mimic the teacher’s dynamics, achieving performance close to the teacher while drastically reducing inference steps. Experiments on LIBERO, RoboTwin 2.0, and real‑robot deployments show significant latency reductions and higher success rates compared to existing few‑step distillation baselines.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Jul 30

Teaching Tiny VLA Models Where to Look and How to Move

arXiv:2607. 04171v3 Announce Type: replace-cross Abstract: Tiny Vision-Language-Action models are appealing for real-time robotic control, but reducing model scale often weakens two capabilities essential for manipulation: task-conditioned spatial grounding and coherent action generation.

By Iok Tong Lei, Ying Jie Yap, Wei Huang, Qingchen Xie, Qianzhi Li, Yujie Zhang, Xiaolong Liu, Zhidong Deng
arXiv Machine Learning
4d ago

XS-VLA: Teaching Tiny Vision-Language-Action Models with Spatial Supervision and Demonstration Conditioning

arXiv:2607.04171v4 Announce Type: replace-cross Abstract: How can richer training supervision improve robot control while keeping the deployed policy compact? We present XS-VLA, a staged training fra...

By Iok Tong Lei, Ying Jie Yap, Wei Huang, Qingchen Xie, Qianzhi Li, Yujie Zhang, Xiaolong Liu, Zhidong Deng
arXiv Machine Learning
Jul 7

XS-VLA: Coupling Coarse-grained Spatial Distillation with Latent Flow Matching for Lightweight Robotic Control

arXiv:2607. 04171v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) have shown strong multimodal understanding and spatial grounding, but their computational cost limits real-time robotic control.

By Lei Iok Tong, Qingchen Xie, Wei Huang, Ying Jie Yap, Yujie Zhang, Qianzhi Li, Xiaolong Liu, Zhidong Deng