arXiv AI

CrossVLA: Cross-Paradigm Post-Training and Inference Optimization for Vision-Language-Action Models

arXiv:2605. 21854v2 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models have rapidly converged on a small set of architectural patterns: discrete-token autoregression (e.

arXiv Computer Vision
Sep 15

What Makes an Efficient VLA? Navigating Action-Head Design, Scaling, and Latency

arXiv:2609.13984v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models combine a pretrained vision encoder, a language backbone, and an action head, but their relative contribution has...

By Luoyang Sun, Guoyang Xia, Fengfa Li, Lei Ren, Xinyu Cui, Haifeng Zhang, Fangxiang Feng, Kaike Zhang, Kun Zhan, Yan Xie, Jun Wang, Cheng Deng
arXiv Computer Vision
Sep 11

IMLE-VLA: Fast Single-Step Action Generation for Vision-Language-Action Policies

IMLE‑VLA replaces the iterative action head in vision‑language‑action policies with a single‑step conditional generator trained via conditional Implicit Maximum Likelihood Estimation (cIMLE). This eliminates multi‑step sampling, boosting inference frequency by 3.67× (55 Hz vs. 15 Hz) and achieving the highest average success rate (98.0 %) on the 40‑task LIBERO benchmark while maintaining robustness under perturbations. Real‑world tests on a Franka Emika Panda show smoother, faster motions and a 3.9×–6.6× reduction in inference time per episode.

By Kian Hosseinkhani (Simon Fraser University), Qinhe Peng (University of Pennsylvania), George Shramko (Simon Fraser University), Mehran Aghabozorgi (Simon Fraser University), Jianing Qian (University of Pennsylvania), Tristan Engst (Simon Fraser University), Alireza Moazeni (Simon Fraser University), Dinesh Jayaraman (University of Pennsylvania), Ke Li (Simon Fraser University, Canada CIFAR AI Chair)
arXiv Computer Vision
Aug 24

Annotations as Rollouts: Efficient and Scalable Reinforcement Learning for Video MLLMs

The paper introduces OraRL, a reinforcement learning framework that leverages annotations as oracle rollouts to improve sample efficiency and scalability for video multimodal large language models (MLLMs). By decoupling advantage estimation and employing sign‑balanced pruning, OraRL achieves faster training and better performance across multiple video‑perception benchmarks compared to existing methods. The approach scales from 0.8B to 9B parameters and handles up to 100k prompts, delivering significant gains in temporal mIoU, tracking accuracy, segmentation, and spatial‑intelligence metrics.

By Yunheng Li, Guohong Mu, Hao Li, Shengsheng Qian, Dingwen Zhang, Qibin Hou, Ming-Ming Cheng
Hugging Face Trending Papers
Jul 29

TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM

Vision-language-action (VLA) models commonly adopt an LLM-centric $V \to L \to A$ pathway, where visual observations are projected into the representation space of a large language model before being decoded into robot actions. Although effective, this design incurs substantial computation and memory overhead at every policy invocation.

arXiv AI
Jul 29

CoTinyVLA: Chain-of-Thought Distillation for a Sub-Billion-Parameter Vision-Language-Action Model

arXiv:2607. 25487v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models translate natural-language commands into robot action sequences, but leading systems on the LIBERO-Plus robustness benchmark use three- to seven-billion-parameter backbones whose memory demands can exceed embedded robotic budgets.

By Minhyeok Lee, Chiyoung Kim, Chanhoe Gu, Seongrok Kim, Sanghyuk Roy Choi, Donghwan Hwang, Donghun Ryu, Seokhyun Kim
arXiv AI
Jun 2

VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models

arXiv:2601. 03309v2 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models, which integrate pretrained large Vision-Language Models (VLM) into their policy backbone, are gaining significant attention for their promising generalization capabilities.

By Jianke Zhang, Xiaoyu Chen, Qiuyue Wang, Mingsheng Li, Yanjiang Guo, Yucheng Hu, Jiajun Zhang, Shuai Bai, Junyang Lin, Jianyu Chen
arXiv Machine Learning
Sep 25

Decoupled Early Exits for Task-Dependent Compute Allocation in Flow-Matching VLAs

The paper introduces a framework for Flow‑Matching Vision‑Language‑Action (VLA) models that allows independent adjustment of backbone depth, action expert depth, and denoising steps. Lightweight Exit Transformers are added at intermediate layers to enable early exits, and a KV Cache synthesis mechanism manages skipped layers so the action expert can exit deeper than the backbone. Experiments on SmolVLA and π0.5 across LIBERO and Meta‑World show that joint tuning of these compute axes reduces latency by 79.2 % and FLOPs by 31.8 %, while improving mean success rate by 5.6 %.

By Riccardo Andrea Izzo, Rimvydas Rubavicius, Gianluca Bardaro, Subramanian Ramamoorthy, Matteo Matteucci, Alessandro Suglia