IMLE-VLA: Fast Single-Step Action Generation for Vision-Language-Action Policies
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2608.29208v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models, built upon Vision-Language Models (VLMs), have significantly enhanced robotic capabilities by leveraging interne...
arXiv:2605. 21854v2 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models have rapidly converged on a small set of architectural patterns: discrete-token autoregression (e.
CF‑VLA introduces a two‑stage coarse‑to‑fine approach for vision‑language‑action policies, replacing multi‑step sampling with a coarse initialization that constructs an action‑aware starting point and a single‑step refinement that corrects residual errors. The coarse stage learns a conditional posterior over endpoint velocity to transform Gaussian noise into a structured initialization, while the fine stage performs a fixed‑time refinement. Experiments on CALVIN and LIBERO demonstrate that CF‑VLA achieves a strong efficiency‑performance trade‑off, reducing action sampling latency by 75.4 % and achieving an 83.0 % real‑robot success rate, outperforming existing NFE=2 methods and matching or surpassing NFE=10 baselines.
arXiv:2608. 11739v1 Announce Type: cross Abstract: The prevailing recipe for Vision-Language-Action (VLA) models couples a pretrained VLM with a separately trained flow-matching action expert.
arXiv:2606. 05737v1 Announce Type: cross Abstract: Diffusion-based vision-language-action (VLA) models often inherit the image-generation view: actions are generated by iterative denoising.
LoopVLA introduces a recurrent Vision‑Language‑Action architecture that learns to refine multimodal representations, predict actions, and estimate when further refinement is unnecessary. By iteratively applying a shared Transformer block and producing a sufficiency score at each step, it decouples refinement from fixed layer indices and aligns confidence scores with action quality through a self‑supervised objective. Experiments on LIBERO, LIBERO‑Plus, and VLA‑Arena demonstrate that LoopVLA reduces model parameters by 45% and boosts inference throughput up to 1.7× while matching or surpassing strong baselines in task success.