arXiv AI

The Speedup Paradox: Rethinking Inference Speed-Quality Trade-off in Embodied Tasks

arXiv:2606. 28529v1 Announce Type: cross Abstract: Embodied foundation models have recently been widely used to improve robot generalization and task success rates.

arXiv AI
Jun 11

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models

arXiv:2606. 11324v1 Announce Type: cross Abstract: We introduce Embodied-R1.

By Yifu Yuan, Yaoting Huang, Xianze Yao, Yutong Li, Shuoheng Zhang, Linqi Han, Pengyi Li, Jiangeng Sun, Wenting Jia, Zhao Zhang, Yuhao Liu, Ruihao Liao, Yucheng Hu, Qiyu Wu, Yuxiao Li, Zibin Dong, Fei Ni, Yan Zheng, Shuyang Gu, Yi Ma, Hongyao Tang, Han Hu, Jianye Hao
arXiv AI
Sep 17

rMuscle: Robotic Muscle Memory for Efficient Vision-Language-Action Model Inference

The paper introduces rMuscle, a real‑time Vision‑Language‑Action inference framework that mimics human muscle memory to accelerate robotic decision making. By exploiting repeated task similarity, rMuscle uses a dual‑phase cache: a Context Cache reuses visual‑token outputs and an Action Cache reuses neuron activation patterns, reducing computation and weight accesses. Experiments on RTX 4090 and Jetson Thor show 1.29–1.42× speedups on LIBERO, RoboTwin, and physical manipulation tasks while preserving success rates on real robots.

By Kaijun Zhou, Zhiyang Li, Le Chen, Jinyu Gu
arXiv Computer Vision
Aug 27

Training-Free Interaction-Aligned Visual Token Pruning for Efficient Embodied Manipulation

The paper introduces Interaction‑Aligned Pruning (IAprune), a training‑free method for visual token pruning in embodied manipulation tasks. IAprune jointly decides per‑frame budget and token selection, using semantic‑motion spatial agreement to choose between conservative and aggressive coverage, and applies geometric residual correction to focus on under‑represented boundaries. Experiments on four policies, three simulation benchmarks, and a real‑robot platform show that IAprune matches unpruned performance on LIBERO while achieving up to 1.54× speed‑up and 1.48× acceleration on a real robot.

By Jintao Cheng, Weibin Li, Haozhe Wang, Gang Wang, Yipu Zhang, Xiaoyu Tang, Jin Wu, Xieyuanli Chen, Yunhui Liu, Wei Zhang
arXiv Computation and Language
4d ago

ProgressCompass: Embodied Progress Reward Models Are Lost Without the Right Context

The paper introduces ProgressCompass, a framework that enhances Embodied Progress Reward Models (PRMs) by providing the necessary contextual information for accurate progress estimation in long manipulation tasks. It presents ContextProgress-Bench, a benchmark with 24 tasks that tests PRMs under three context-dependent scenarios—State Recall, Sequence Tracking, and Recurrence Disambiguation—showing that even history-aware PRMs struggle without proper context. By integrating a context-aware loop that leverages general-purpose vision‑language models, ProgressCompass reduces PRM progress error by up to 82% and improves rank agreement by 76%.

By Jianshu Zhang, Keliang Wu, Chengxuan Qian, Xiyuan Yang, Ce Zhang, Ariel Tian, Anbang Liu, Haoran Lu, Han Liu
arXiv AI
Jul 7

Cortex: A Bidirectionally Aligned Embodied Agent Framework for Long-horizon Manipulation

arXiv:2607. 05377v1 Announce Type: cross Abstract: While recent Vision-Language-Action (VLA) models show promise toward generalist manipulation policies, they struggle with long-horizon tasks due to their Markovian nature-relying solely on current observations.

By Jiaqi Peng, Xiqian Yu, Delin Feng, Yuqiang Yang, Wenzhe Cai, Jing Xiong, Ganlin Yang, Jinliang Zheng, Jiafei Cao, Xueyuan Wei, Jiangmiao Pang, Yuan Shen, Tai Wang