arXiv:2606. 08094v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) policies are typically shipped as Python/PyTorch stacks that assume a workstation-class GPU, a mismatch for the hardware on which robots actually run.
By Khanh D. Nguyen, Hung T. Ho, Chinh T. Nguyen, Thanh Q. Duong, Linh D. Le, Duy M. H. Nguyen, Vien A. Ngo, An T. Le
arXiv:2608. 12932v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models promise to bring end-to-end reasoning to autonomous driving, but their computational cost remains far too high for real-time control.
By Zekai Li, Yihao Liang, Hongfei Zhang, Jian Chen, Yesheng Liang, Zhijian Liu
Vision-Language-Action (VLA) models promise to bring end-to-end reasoning to autonomous driving, but their computational cost remains far too high for real-time control. The core challenge is structural: VLA inference is not a single bottleneck but a cascade of four.
Robion is a new serving and management system designed to run Vision‑Language‑Action (VLA) models on multi‑GPU edge servers for robot factories. It splits the VLM and ADiT stages within a single GPU, shares streams across multiple models, and prioritizes requests by remaining SLO time, enabling high robot load while meeting strict latency requirements. In experiments, Robion achieves 6.7× higher robot load than vLLM‑Omni and 1.5× higher than a monolithic pipeline, and can serve 64 robots on a 4‑GPU server with 98% SLO attainment.
By Dionysios Adamopoulos, Nattapol Chanpaisit, Basel Fakhri, Christina Giannoula
arXiv:2607. 10942v1 Announce Type: cross Abstract: Physical AI systems, such as autonomous vehicles and intelligent machines, require transformer-based perception models that satisfy stringent edge latency and energy constraints.
By Ashiyana Abdul Majeed, Mahmoud Meribout, Neethu Joseph, Abel Kidane Haile, Mohammad Abdullah Al Faruque
arXiv:2607. 12659v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have achieved impressive performance on diverse embodied tasks.
By Zebin Yang, Qi Wang, Yunhe Wang, Xiurui Guo, Bo Yu, Shaoshan Liu, Jiafeng Xu, Hao Dong, Meng Li
Visual Perception Engine (VPEngine) is a modular framework that enables efficient GPU usage for robotic vision tasks by sharing a foundation model backbone across multiple specialized task heads. It eliminates redundant feature extraction, supports dynamic task prioritization, and achieves up to 3× speedup over sequential execution. The open‑source Python implementation, with ROS2 C++ bindings, delivers real‑time performance (≥50 Hz) on NVIDIA Jetson Orin AGX using TensorRT‑optimized models.
By Jakub {\L}ucki, Jonathan Becktor, Georgios Georgakis, Rob Royce, Shehryar Khattak
arXiv:2608. 01563v1 Announce Type: new Abstract: Training and deployed inference often cross export, conversion, and platform-specific runtime boundaries.
By Dzmitry Malyshau
arXiv:2606.08684v2 Announce Type: replace
Abstract: We present BLUE, a minimal method for better language use in vision-language-action (VLA) models for autonomous driving (AD). Through extensive ana...
By George Ling, Lijin Yang, Hao Yang, Zhongzhan Huang
SlackDrive is a pre‑inference compute allocator that dynamically selects the compute budget for each driving control step by reusing the realized latency from previous inferences. By profiling a small set of discrete budgets once, it estimates the current compute state online and chooses the highest‑utility budget that stays within the admissible latency envelope. On the NAVSIM v2 benchmark with DriveDreamer‑Policy, SlackDrive boosts latency‑constrained EPDMS performance by 21.7% compared to the best baseline, while full‑budget and token‑pruning approaches exceed the latency limits under runtime contention.
By Xiaohuan Pei, Hengguang Zhou, Yuanhao Ban, Justin Cui, Jiaqi Feng, Haoyu Xie, Tao Huang, Pichao Wang, Yanchao Yang, Cho-Jui Hsieh
arXiv:2609.18623v1 Announce Type: new
Abstract: State-of-the-art vision-language-action models (VLA) for autonomous driving face critical limitations: excessive parameter counts, inefficient high-res...
By Kemal Oksuz, Alexandru Buburuzan, Yuhan Yao, Puneet K. Dokania
Vision-language-action (VLA) models commonly adopt an LLM-centric $V \to L \to A$ pathway, where visual observations are projected into the representation space of a large language model before being decoded into robot actions. Although effective, this design incurs substantial computation and memory overhead at every policy invocation.