arXiv:2609.34911v2 Announce Type: replace-cross
Abstract: Modern robot policies predict a chunk of future actions from a single observation, execute only a prefix, and discard the rest before replann...
By Taesung Kwon, Jangho Park, Sunwoo Park, Youngmin Kim, Seonghyun Jin, Youngjun Jun, Kyumin Choi, Jong Chul Ye
The paper introduces a framework for Flow‑Matching Vision‑Language‑Action (VLA) models that allows independent adjustment of backbone depth, action expert depth, and denoising steps. Lightweight Exit Transformers are added at intermediate layers to enable early exits, and a KV Cache synthesis mechanism manages skipped layers so the action expert can exit deeper than the backbone. Experiments on SmolVLA and π0.5 across LIBERO and Meta‑World show that joint tuning of these compute axes reduces latency by 79.2 % and FLOPs by 31.8 %, while improving mean success rate by 5.6 %.
By Riccardo Andrea Izzo, Rimvydas Rubavicius, Gianluca Bardaro, Subramanian Ramamoorthy, Matteo Matteucci, Alessandro Suglia
arXiv:2609.36540v1 Announce Type: cross
Abstract: Generalist robot policies such as vision-language-action models (VLAs) have achieved remarkable generalization, but their inference delays can confli...
By Moritz Zoellner, Reece O'Mahoney, Ioannis Havoutis, Rohan Paleja
The paper investigates how small action errors evolve when using action chunking in behavioural cloning. By injecting errors at each state and observing their growth under open‑loop (no replanning) and closed‑loop (replanning) regimes, the authors classify states as contracting, expanding, or unresolved. Across twelve manipulation tasks, they find that stable states are rare, error amplification is common, and that short‑horizon fitting can overestimate long‑horizon propagation. Predictors trained on camera and proprioceptive data can recover open‑loop stability but only partially capture closed‑loop dynamics, indicating that standard imitation learning does not reliably produce policies that contract errors when perturbed.
By Aryan Goyal
arXiv:2607. 15621v1 Announce Type: cross Abstract: Large language models bring instruction following and scene reasoning to end-to-end driving, but their inference latency collides with the control rate a vehicle requires.
By Yun Li, Jiachen Gong, Simon Thompson, Ehsan Javanmardi, Qunli Zhang, Zifan Zeng, Shiming Liu, Peng Wang, Zixuan Guo, Manabu Tsukada
The paper argues that diffusion-based action policies can use a frozen, observation‑free backbone as a reusable trajectory prior, with task adaptation handled entirely by the conditioning pathway. By pretraining a general action head on forward‑kinematics data and then freezing it, the authors show that a single backbone can match or outperform normally trained models on MimicGen and LIBERO. Their experiments reveal that a small 5 M‑parameter MLP backbone can rival large U‑Net and transformer backbones, indicating that action backbones are often over‑parameterized and that image‑style architectures may not be the best fit for low‑dimensional action generation.
By Jian Zhou, Sihao Lin, Shuai Fu, Zerui Li, Gengze Zhou, Qi WU