Flow matching (FM) has become a popular action head paradigm for modern embodied models. However, as a conditional generative model, it does not explicitly expose its inherent uncertainty, producing faulty action chunks even when it misinterprets the scene or encounters out-of-distribution (OOD) inputs.
arXiv:2607. 27933v3 Announce Type: replace Abstract: Flow matching (FM) has become a popular action head paradigm for modern embodied models.
By Ziyang Rao, Yiren Zhao, Weiyu Guo, Ben Fei, Yandong Guo, Hui Xiong
arXiv:2606. 18043v1 Announce Type: cross Abstract: Vision-language-action models (VLAs) combine vision-language backbones with expressive generative action heads trained via flow matching on large-scale robotic datasets.
By Ralf R\"omer, Maximilian Seeliger, Saida Liu, Ben Sturgis, Marco Bagatella, Daniel Marta, Andreas Krause, Angela P. Schoellig
arXiv:2605. 00941v4 Announce Type: replace Abstract: Flow matching has become a leading framework for generative modeling, but quantifying the uncertainty of its samples remains an open problem.
By Jiarui Xing, Song Wang, Jian Wang
Vision-language-action models (VLAs) combine vision-language backbones with expressive generative action heads trained via flow matching on large-scale robotic datasets. Despite their strong empirical performance in robotic manipulation, VLAs lack mechanisms to quantify confidence in their predictions and to detect when their actions may be unreliable.
GeoAAC introduces a geometry-based adaptive action chunking technique for Vision‑Language‑Action policies, dynamically adjusting the action horizon based on the reliability of current action predictions. By leveraging the geometric variation in Flow Matching denoising trajectories, GeoAAC constructs a horizon‑wise geometric profile that determines the action horizon during a single generation without extra training. Experiments on LIBERO, LIBERO‑Pro, RoboCasa365, and real‑world manipulation tasks demonstrate consistent gains over fixed‑horizon baselines, achieving up to 8.7 percentage points improvement in simulation and raising real‑world success rates from 53.3% to 74.4%.
By Xin Chen, Sen Chen, Yujuan Ding, Jian Liu, Guoqing Wang, Wei Ye, Heng Tao Shen, Yi Bin
arXiv:2607. 29235v1 Announce Type: cross Abstract: Although world-action models (WAMs) enhance long-horizon robot control by predicting visual evolution before acting, long-horizon reliability demands repeated re-grounding in real observations--not recursive rollout.
By Peize Li, Ruimeng Zhang, Ru Zhang, Cong Huang, Kai Chen, Shanghang Zhang
arXiv:2604. 18194v2 Announce Type: replace Abstract: Single-step generators promise high-fidelity synthesis at a fraction of the inference and training cost of ordinary differential equation (ODE)-based flow models, a central concern when compute is limited.
By Arkadii Kazanskii, Tatiana Petrova, Andrey Ustyuzhanin, Konstantin Bagrianskii, Aleksandr Puzikov, Radu State
arXiv:2610.01193v1 Announce Type: cross
Abstract: Counterfactual generation seeks to sample outcomes under a hypothetical intervention or decision using observational data collected under the factual...
By Yunrui Guan, Krishnakumar Balasubramanian, Shiva Prasad Kasiviswanathan
CF‑VLA introduces a two‑stage coarse‑to‑fine approach for vision‑language‑action policies, replacing multi‑step sampling with a coarse initialization that constructs an action‑aware starting point and a single‑step refinement that corrects residual errors. The coarse stage learns a conditional posterior over endpoint velocity to transform Gaussian noise into a structured initialization, while the fine stage performs a fixed‑time refinement. Experiments on CALVIN and LIBERO demonstrate that CF‑VLA achieves a strong efficiency‑performance trade‑off, reducing action sampling latency by 75.4 % and achieving an 83.0 % real‑robot success rate, outperforming existing NFE=2 methods and matching or surpassing NFE=10 baselines.
By Fan Du, Feng Yan, Jianxiong Wu, Xinrun Xu, Weiye Zhang, Weinong Wang, Yu Guo, Bin Qian, Zhihai He, Fei Wang, Heng Yang
The paper introduces Kinematic MeanFlow (K-MF), a one‑step action generation policy for Robotic Foundation Models that addresses instability in the MeanFlow framework. By decoupling the time derivative into two sub‑interval terms, K-MF captures early and late denoising dynamics separately, reducing error amplification. Experiments show K-MF achieves faster inference—reducing action‑head latency by 67.5%–74.4% and overall end‑to‑end latency by 30.3%–54.9%—while outperforming multi‑step flow matching on various tasks.
By Jiawei Fan, Sifeng Wang, Yuqing Hou, Anbang Yao
arXiv:2606. 05254v1 Announce Type: new Abstract: World-action models (WAMs) jointly generate future video and robot actions through iterative diffusion, achieving strong performance on manipulation benchmarks but requiring tens of denoising steps, a cost that precludes real-time control.
By Arman Akbari, Ci Zhang, Arash Akbari, Lin Zhao, Yixiao Chen, Weiwei Chen, Xuan Zhang, Geng Yuan, Yanzhi Wang