arXiv:2606. 05737v1 Announce Type: cross Abstract: Diffusion-based vision-language-action (VLA) models often inherit the image-generation view: actions are generated by iterative denoising.
By Yitong Chen, Shiduo Zhang, Jingjing Gong, Xipeng Qiu
The paper argues that diffusion-based action policies can use a frozen, observation‑free backbone as a reusable trajectory prior, with task adaptation handled entirely by the conditioning pathway. By pretraining a general action head on forward‑kinematics data and then freezing it, the authors show that a single backbone can match or outperform normally trained models on MimicGen and LIBERO. Their experiments reveal that a small 5 M‑parameter MLP backbone can rival large U‑Net and transformer backbones, indicating that action backbones are often over‑parameterized and that image‑style architectures may not be the best fit for low‑dimensional action generation.
By Jian Zhou, Sihao Lin, Shuai Fu, Zerui Li, Gengze Zhou, Qi WU
Dynin‑Robotics introduces an omnimodal masked‑diffusion backbone, Dynin‑Omni, that jointly represents language, visual observations, goals, and actions as discrete tokens. By conditioning on different spans, the same model learns action prediction, next‑observation prediction, goal‑state prediction, and trajectory‑to‑instruction reconstruction, enabling test‑time scaling through goal prediction and action‑candidate evaluation. The system, pretrained on 1.33 million trajectories from 48 Open X‑Embodiment datasets, achieves competitive performance on LIBERO, zero‑shot LIBERO‑Plus, and a 78.4 % success rate on a Franka Research 3 robot, while a block‑parallel implementation speeds up action decoding by up to 29.2×.
By Hoeun Lee, Jaeik Kim, Jusang Oh, Jinhyeok Kim, Geon Choi, Hyeonggeun Kim, Jaeyoung Do
arXiv:2508. 20072v4 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models adapt large vision-language backbones to map images and instructions into robot actions.
By Zhixuan Liang, Yizhuo Li, Tianshuo Yang, Chengyue Wu, Sitong Mao, Liuao Pei, Tian Nian, Shunbo Zhou, Xiaokang Yang, Jiangmiao Pang, Yao Mu, Ping Luo
arXiv:2607. 27138v1 Announce Type: cross Abstract: Vision-language-action (VLA) models remain constrained by scarce action-labeled robot data, whereas action-free videos offer abundant observations of physical change.
By Zuojin Tang, Feifan Luo, Haoyun Liu, Botai Yuan, Dekang Qi, Ronghan Chen, Yandan Yang, Tong Lin, Xinyuan Chang, Mu Xu, Bin Liu, De Ma, Zhiheng Ma
arXiv:2610.02666v1 Announce Type: new
Abstract: Vision-Language-Action (VLA) models map visual observations and language instructions to continuous robot actions, but a diffusion-based action expert...
By Jin Hyun, Jung Gyu Min, Gyuhyun Jung, Youngjoo Lee
arXiv:2609.36540v1 Announce Type: cross
Abstract: Generalist robot policies such as vision-language-action models (VLAs) have achieved remarkable generalization, but their inference delays can confli...
By Moritz Zoellner, Reece O'Mahoney, Ioannis Havoutis, Rohan Paleja
arXiv:2607. 19919v1 Announce Type: cross Abstract: We propose Diffusion ReRoll, a diffusion-based framework for robotic sequential prediction that enables revisable denoising over horizons.
By Seonsoo Kim, Seongil Hong, Jun-Gill Kang
DriftOPD is a teacher‑free, rollout‑free framework that performs sequence‑level on‑policy distillation of continuous Vision‑Language‑Action (VLA) action experts. It decomposes the sequence‑level reverse‑KL divergence into a chunk‑level reverse‑KL term and a future‑potential term, optimizing them with a one‑step drifting objective and a Q‑function critic learned from offline demonstrations. Experiments on multiple VLA architectures in simulation and real‑world manipulation show that DriftOPD outperforms existing one‑step distillation baselines while matching the task success of multi‑step teacher policies.
By Youngjun Jun, Kyumin Choi, Youngmin Kim, Seonghyun Jin, Sunwoo Park, Jangho Park, Jong Chul Ye
arXiv:2606. 08657v1 Announce Type: cross Abstract: Diffusion-based visuomotor policies operating directly in raw action spaces conflate scene comprehension with trajectory generation within a single denoising process.
By Zhexuan Zhou, Yichen Lai, Jinhao Zhang, Huizhe Li, Youmin Gong, Jie Mei
arXiv:2607. 09818v1 Announce Type: cross Abstract: Vision-language-action (VLA) models aim to understand natural-language instructions and visual observations, and to generate and execute corresponding actions as embodied agents.
By Shengzhuo Yang, Ronghao Yu, Chuanjie Lv, Linpeng Peng, Hang Yu, Jie Ren, Jiajun Lv, Yong Liu
The paper introduces Intention Distillation (INDI), a method that injects behavior-level intent into Vision‑Language‑Action (VLA) model decoders by leveraging a frozen teacher vision‑language model to interpret demonstrations. During training, the teacher processes the current observation, instruction, coarse action summary, and execution video, producing a multimodal intent representation that the VLA decoder uses alongside trajectory and execution features to predict actions. Experiments on SimplerEnv‑Bridge, RoboCasa Kitchen, and real‑world tasks show that INDI consistently improves success rates, especially on longer‑horizon tasks, demonstrating that explicit modeling of semantic intent benefits action decoders.
By Sangoh Lee, Sangwoo Mo, Wook-Shin Han