arXiv AI

The Geometry of Flow-Matching Uncertainty: A Cost-free Uncertainty Proxy and Its Application in Flow-based VLA Failure Detection

arXiv:2607. 27933v3 Announce Type: replace Abstract: Flow matching (FM) has become a popular action head paradigm for modern embodied models.

arXiv Machine Learning
Jun 17

Uncertainty Quantification for Flow-Based Vision-Language-Action Models

arXiv:2606. 18043v1 Announce Type: cross Abstract: Vision-language-action models (VLAs) combine vision-language backbones with expressive generative action heads trained via flow matching on large-scale robotic datasets.

By Ralf R\"omer, Maximilian Seeliger, Saida Liu, Ben Sturgis, Marco Bagatella, Daniel Marta, Andreas Krause, Angela P. Schoellig
Hugging Face Trending Papers
Jun 16

Uncertainty Quantification for Flow-Based Vision-Language-Action Models

Vision-language-action models (VLAs) combine vision-language backbones with expressive generative action heads trained via flow matching on large-scale robotic datasets. Despite their strong empirical performance in robotic manipulation, VLAs lack mechanisms to quantify confidence in their predictions and to detect when their actions may be unreliable.

arXiv AI
Sep 18

GeoAAC: Geometry-Based Adaptive Action Chunking from Denoising Trajectories in VLA Policies

GeoAAC introduces a geometry-based adaptive action chunking technique for Vision‑Language‑Action policies, dynamically adjusting the action horizon based on the reliability of current action predictions. By leveraging the geometric variation in Flow Matching denoising trajectories, GeoAAC constructs a horizon‑wise geometric profile that determines the action horizon during a single generation without extra training. Experiments on LIBERO, LIBERO‑Pro, RoboCasa365, and real‑world manipulation tasks demonstrate consistent gains over fixed‑horizon baselines, achieving up to 8.7 percentage points improvement in simulation and raising real‑world success rates from 53.3% to 74.4%.

By Xin Chen, Sen Chen, Yujuan Ding, Jian Liu, Guoqing Wang, Wei Ye, Heng Tao Shen, Yi Bin
arXiv AI
Jul 21

RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations

arXiv:2510. 00037v5 Announce Type: replace-cross Abstract: In Vision-Language-Actionf(VLA) models, robustness to real-world perturbations is critical for deployment.

By Jianing Guo, Zhenhong Wu, Chang Tu, Yiyao Ma, Xiangqi Kong, Zhiqian Liu, Jiaming Ji, Shuning Zhang, Yuanpei Chen, Kai Chen, Qi Dou, Yaodong Yang, Xianglong Liu, Huijie Zhao, Weifeng Lv, Simin Li
arXiv AI
Aug 3

FBFM: A Training-Free Asynchronous Feedback Mechanism for Flow-Matching in World-Action Models Execution

arXiv:2607. 29235v1 Announce Type: cross Abstract: Although world-action models (WAMs) enhance long-horizon robot control by predicting visual evolution before acting, long-horizon reliability demands repeated re-grounding in real observations--not recursive rollout.

By Peize Li, Ruimeng Zhang, Ru Zhang, Cong Huang, Kai Chen, Shanghang Zhang
arXiv Machine Learning
Sep 21

$\lambda$-Controlled GRPO: Turning Flow-Matching Ratio Instability into a Budgeted Resource

arXiv:2609. 22041v1 Announce Type: new Abstract: Reinforcement learning is increasingly used to align image generators with reward signals, and Flow-GRPO recently extended this paradigm to flow-matching models by treating the denoising sampler as a stochastic policy that can be optimized from reward feedback.

By Yufeng Wang, Parivesh Priye, Meeshawn Marathe, Ramit Pahwa
arXiv Machine Learning
Sep 25

Decoupled Early Exits for Task-Dependent Compute Allocation in Flow-Matching VLAs

The paper introduces a framework for Flow‑Matching Vision‑Language‑Action (VLA) models that allows independent adjustment of backbone depth, action expert depth, and denoising steps. Lightweight Exit Transformers are added at intermediate layers to enable early exits, and a KV Cache synthesis mechanism manages skipped layers so the action expert can exit deeper than the backbone. Experiments on SmolVLA and π0.5 across LIBERO and Meta‑World show that joint tuning of these compute axes reduces latency by 79.2 % and FLOPs by 31.8 %, while improving mean success rate by 5.6 %.

By Riccardo Andrea Izzo, Rimvydas Rubavicius, Gianluca Bardaro, Subramanian Ramamoorthy, Matteo Matteucci, Alessandro Suglia
arXiv AI
2d ago

Kinematic MeanFlow: One-Step Action Generation Policy for Robotic Foundation Models

The paper introduces Kinematic MeanFlow (K-MF), a one‑step action generation policy for Robotic Foundation Models that addresses instability in the MeanFlow framework. By decoupling the time derivative into two sub‑interval terms, K-MF captures early and late denoising dynamics separately, reducing error amplification. Experiments show K-MF achieves faster inference—reducing action‑head latency by 67.5%–74.4% and overall end‑to‑end latency by 30.3%–54.9%—while outperforming multi‑step flow matching on various tasks.

By Jiawei Fan, Sifeng Wang, Yuqing Hou, Anbang Yao
arXiv Machine Learning
Jun 5

Flash-WAM: Modality-Aware Distillation for World Action Models

arXiv:2606. 05254v1 Announce Type: new Abstract: World-action models (WAMs) jointly generate future video and robot actions through iterative diffusion, achieving strong performance on manipulation benchmarks but requiring tens of denoising steps, a cost that precludes real-time control.

By Arman Akbari, Ci Zhang, Arash Akbari, Lin Zhao, Yixiao Chen, Weiwei Chen, Xuan Zhang, Geng Yuan, Yanzhi Wang