arXiv Machine Learning

Latent Telepathy: Multi-Robot Communication with Self-Supervised Perceptual Latents

arXiv AI
4d ago

Cooperative Multi-Agent Vision-Language-Action Models via Reinforced Fine Tuning

arXiv:2609.36588v1 Announce Type: cross Abstract: We study reinforcement learning (RL) methods for cooperative multi-agent Vision-Language-Action (VLA) models. This problem is challenging because VLA...

By Ruixiao Xu, Wong Lik Hang Kenny, Zhiqian Liu, Jianing Guo, Hanxiao Li, Kejian Shi, Shuning Zhang, Pu Feng, Yongjia Ma, Yuqing Ma, Kai Chen, Qi Dou, Yaodong Yang, Xianglong Liu, Simin Li
arXiv AI
Aug 25

SRMT: Shared Memory for Multi-agent Lifelong Pathfinding

The paper introduces the Shared Recurrent Memory Transformer (SRMT), a decentralized multi‑agent reinforcement learning framework that uses a global memory workspace for agents to broadcast and query each other’s learned states. SRMT is evaluated on the Partially Observable Multi‑Agent Pathfinding (PO‑MAPF) problem, showing that shared memory enables emergent coordination even with minimal reward guidance and outperforms existing baselines on the Bottleneck task and scales competitively on larger POGEMA maps. The authors provide open‑source code for training and evaluation on GitHub.

By Alsu Sagirova, Yuri Kuratov, Mikhail Burtsev
arXiv AI
Sep 11

Modality-Decoupled Federated Learning for Privacy-Preserving Embodied Intelligence in 6G

The paper introduces FedMVLA, a modality‑decoupled federated learning framework designed for privacy‑preserving embodied intelligence in 6G networks. It addresses the unique challenges of vision‑language‑action models by employing modality‑aware federated aggregation, privacy allocation, and communication compression, along with a precision‑critical action transport slice. A case study on federated robotic manipulation demonstrates significant gains in task success, scalability, and uplink payload reduction compared to standard FedAvg.

By Zhuodong Liu, Xiangyu Li, Chunhong Yuan, Hongyang Du, Bodong Shang, Qingqing Wu, Tony Q. S. Quek, Mohsen Guizani
arXiv AI
Aug 13

G0.5: One Autoregressive Stream for Robot Reasoning and Action

arXiv:2608. 11739v1 Announce Type: cross Abstract: The prevailing recipe for Vision-Language-Action (VLA) models couples a pretrained VLM with a separately trained flow-matching action expert.

By Yicheng Liu, Zibin Dong, Baijun Ye, Tianyuan Yuan, Tao Jiang, Anqi Yang, Shicheng Cao, Haonan Liu, Yue Sun, Zihan Guo, Xiao Liu, Dong Ke, Changxun Pan, Chenru Wu, Tailai Cheng, Xiaoshu Ren, Xinlei Zhang, Jianning Cui, Zijie Zhao, Haoyu Zhang, Kaiming Xu, Haodong Yang, Bowen Zhang, Jiahui Niu, Shaoting Zhu, Shiduo Zhang, Hang Zhao