arXiv AI

Workspace Models: Lightweight Robotic Memory via Saliency-Driven Supervision

The paper introduces the workspace token, a lightweight latent memory representation for robotic manipulation that captures task-relevant historical information. By training with VLM queries only during training, the token can be queried efficiently at deployment, replacing full observations. Experiments in simulation and on hardware show that policies using the workspace token solve memory-intensive tasks without in‑loop VLM reasoning and even outperform heavier approaches.

arXiv AI
Aug 5

Gated Memory Policy: In-Context Memorization and Adaptation

arXiv:2604. 18933v2 Announce Type: replace-cross Abstract: Robotic manipulation tasks exhibit varying memory requirements, ranging from Markovian tasks that require no memory to non-Markovian tasks that demand in-context memorization of historical information within a single trial or in-context adaptation based on the outcomes of multiple past trials.

By Yihuai Gao, Jeff Jinyun Liu, Shuang Li, Shuran Song
arXiv Computer Vision
3d ago

D$^2$-VLA: Dual-Memory Dual-Frequency Vision-Language-Action Model For Long Dynamic Manipulation

arXiv:2609.34792v2 Announce Type: replace Abstract: Long-horizon manipulation requires robots to remember cues that are no longer in view while responding to moving objects. Yet vision-language-actio...

By Zijian Ye, Chengqi Wei, Wei Huang, Anlin Zheng, Chunyu Zou, Liangyu Wu, Zikang Zhao, Zhenjie Peng, Yushuo Yang, Shuman Zhao, Zhongrui Wang, Xiaojuan Qi
arXiv AI
Jul 24

VPWEM: Non-Markovian Visuomotor Policy with Working and Episodic Memory

arXiv:2603. 04910v2 Announce Type: replace-cross Abstract: Imitation learning from human demonstrations has achieved significant success in robotic control, yet most visuomotor policies still condition on single-step observations or short-context histories, making them struggle with non-Markovian tasks that require long-term memory.

By Yuheng Lei, Zhixuan Liang, Hongyuan Zhang, Ping Luo
arXiv AI
Sep 17

rMuscle: Robotic Muscle Memory for Efficient Vision-Language-Action Model Inference

The paper introduces rMuscle, a real‑time Vision‑Language‑Action inference framework that mimics human muscle memory to accelerate robotic decision making. By exploiting repeated task similarity, rMuscle uses a dual‑phase cache: a Context Cache reuses visual‑token outputs and an Action Cache reuses neuron activation patterns, reducing computation and weight accesses. Experiments on RTX 4090 and Jetson Thor show 1.29–1.42× speedups on LIBERO, RoboTwin, and physical manipulation tasks while preserving success rates on real robots.

By Kaijun Zhou, Zhiyang Li, Le Chen, Jinyu Gu
arXiv Computer Vision
Sep 25

Self-Adaptive VLA for Robust Robot Deployment

The paper introduces Self‑Adaptive VLA, a post‑training method that lets Vision‑Language‑Action policies self‑adapt to deployment‑time hardware shifts by using rollouts as context. It creates shift‑conditioned expert demonstrations, compresses visual, proprioceptive, and action data into a latent context token, and modulates the policy via adaptive layer normalization. Experiments on four precision‑critical manipulation tasks show the method recovers over 80 % of the base policy’s performance under actuation bias and encoder offsets, and improves robustness on new workstations.

By Hongxin Zhang, Chunru Lin, Tsun-Hsuan Wang, Zhenjia Xu, Chuang Gan