arXiv:2609.36471v1 Announce Type: cross
Abstract: World-Action Models (WAMs) improve robotic manipulation by conditioning action generation on predicted future observations, but future prediction add...
By Guoheng Sun, Chen Chen, Jin Wang, Ang Li, Teresa Lv
CF‑VLA introduces a two‑stage coarse‑to‑fine approach for vision‑language‑action policies, replacing multi‑step sampling with a coarse initialization that constructs an action‑aware starting point and a single‑step refinement that corrects residual errors. The coarse stage learns a conditional posterior over endpoint velocity to transform Gaussian noise into a structured initialization, while the fine stage performs a fixed‑time refinement. Experiments on CALVIN and LIBERO demonstrate that CF‑VLA achieves a strong efficiency‑performance trade‑off, reducing action sampling latency by 75.4 % and achieving an 83.0 % real‑robot success rate, outperforming existing NFE=2 methods and matching or surpassing NFE=10 baselines.
By Fan Du, Feng Yan, Jianxiong Wu, Xinrun Xu, Weiye Zhang, Weinong Wang, Yu Guo, Bin Qian, Zhihai He, Fei Wang, Heng Yang
arXiv:2607. 02222v1 Announce Type: cross Abstract: Vision-Language Navigation has increasingly emphasized high-level instruction reasoning, memory, global map construction, and instruction decomposition, while the low-level action representation remains comparatively underexplored.
By Haokun Liu, Zhaoqi Ma, Yicheng Chen, Wentao Zhang, Masaki Kitagawa, Zicen Xiong, Jinjie Li, Moju Zhao
arXiv:2607. 15065v1 Announce Type: cross Abstract: Predictive world models enable robots to plan by imagining the outcomes of their actions, but their value for control hinges on generating many rollouts quickly.
By Susie Lu, Haonan Chen, Weirui Ye, Yilun Du
PACT‑WAM is a world‑action model that simultaneously predicts a 16‑step action trajectory and its corresponding visual forecast for robot manipulation. It uses a hierarchical history encoder that compresses past observations into fewer tokens, reducing processing cost by 75% compared to dense encoding. The model’s shared flow module updates action and visual states jointly, and a TiTok‑VAE decoder reconstructs multi‑view future images, which are then used by a vision‑language component (Proposal Review) to improve execution‑prefix selection and proposal rejection, boosting success rates on several benchmarks.
By Yushan Liu, Jingjing Fan, Shoujie Li, Yifan Xie, Xiao-Ping Zhang, Wenbo Ding
DeltaWAM introduces a new approach to world-action models (WAMs) for bimanual manipulation by jointly predicting visual deltas and actions instead of dense future frames, thereby reducing redundant modeling of unchanged content and mitigating nuisance appearance variations. The method employs three architectures with varying representation and computation sharing, and incorporates Streaming Delta Memory (SDM) to update cached anchor context using compact observed deltas, which cuts heavy video-expert processing. Experiments on RoboTwin show that DeltaWAM with SDM raises average success rates from 81.3% to 85.4% in clean settings and from 75.8% to 83.9% under visual randomization, while also reducing training FLOPs by up to 23.77% and inference latency by 36.57%.
whyItMatters":"DeltaWAM improves both performance and computational efficiency for bimanual manipulation tasks by focusing on visual deltas and efficient memory updates, as demonstrated by higher success rates and lower FLOPs on RoboTwin."
By Han Yan, Zishang Xiang, Haokai Jiang, Zeyu Zhang, Qilin Wang, Weiyu Guo, Yandong Guo, Boxin Shi, Hao Tang