arXiv:2608.21407v1 Announce Type: cross
Abstract: Vision-language-action (VLA) models face a crucial tradeoff between their task success rate and the policy-call frequency. Executing a single action...
By Farida Mohsen, Thowayba Elkaffash, Mohammad Reza Chalak Qazani, Mohamed Mabrok, Nader Meskin, Ali Safa
arXiv:2609.26420v1 Announce Type: cross
Abstract: Text-to-motion models generate plausible human motion but do not model a robot's dynamics; whole-body tracking controllers execute robot references r...
By Raphael Memmesheimer, Sven Behnke
arXiv:2606. 29898v1 Announce Type: cross Abstract: Real-world evaluation is the gold standard for robot policies because it tests them against the physical conditions and deployment challenges they are ultimately designed to handle.
By Haoxu Huang, Tongsam Zheng, Yifan Chen, Jiacheng You, Yang Gao
arXiv:2609.36808v1 Announce Type: cross
Abstract: Current embodied models do not respond to their own failures, although what just went wrong could inform a small adjustment on the next attempt, the...
By Long Li, Qichao Zhao, Yue Yang, Fan Xu, Zhe Wang, Alan Wee-Chung Liew, Chao Qu, Heng Tao Shen, Shirui Pan
arXiv:2607. 16821v1 Announce Type: cross Abstract: Task arithmetic, sequential fine-tuning, activation steering, and first-order random search all operate through relatively small perturbations around an already trained checkpoint, and they rely on different local approximations: individual perturbations should be first-order predictable, task updates should compose with controlled interference, useful tangent structure should be stable and possible to estimate, and weight edits should have counterparts in representation space.
By Irina Piontkovskaia, Sergey Nikolenko
arXiv:2610.00604v1 Announce Type: cross
Abstract: Vision-language-action policies often see only one or a few recent frames, which makes it difficult to evaluate how they use information that disappe...
By Egor Cherepanov, Nikita Kachaev, Aleksandr I. Panov, Alexey K. Kovalev