arXiv AI
Aug 13

G0.5: One Autoregressive Stream for Robot Reasoning and Action

arXiv:2608. 11739v1 Announce Type: cross Abstract: The prevailing recipe for Vision-Language-Action (VLA) models couples a pretrained VLM with a separately trained flow-matching action expert.

By Yicheng Liu, Zibin Dong, Baijun Ye, Tianyuan Yuan, Tao Jiang, Anqi Yang, Shicheng Cao, Haonan Liu, Yue Sun, Zihan Guo, Xiao Liu, Dong Ke, Changxun Pan, Chenru Wu, Tailai Cheng, Xiaoshu Ren, Xinlei Zhang, Jianning Cui, Zijie Zhao, Haoyu Zhang, Kaiming Xu, Haodong Yang, Bowen Zhang, Jiahui Niu, Shaoting Zhu, Shiduo Zhang, Hang Zhao
arXiv Computer Vision
Sep 24

Task-Prototype Guided Flow Matching for Few-Shot Generalization in Vision-Language Robot Manipulation

Task-Prototype Guided Flow Matching (TP-Flow) is a few‑shot manipulation framework that transforms support demonstrations into structured task‑prototype tokens to guide both the initial flow prior and the velocity field. It uses symmetric cross‑attention with learnable queries to extract phase‑level prototypes, parameterizes a task‑adaptive initial distribution, and injects prototype information through gated adaptive normalization. TP‑Flow is trained with an episodic support‑query objective and prototype contrastive regularization, achieving high success rates on the LEROBOT‑ARM‑SO101 platform while maintaining real‑time execution and low latency.

By Yizhao Wang, Guantao Zhang, Jingbo Wang
arXiv Machine Learning
Jun 4

Breaking the Scale Barrier: One-Shot Knowledge Transfer via Frequency Transform

arXiv:2603. 07523v3 Announce Type: replace Abstract: Transferring knowledge by fine-tuning large-scale pre-trained networks has become a standard paradigm for downstream tasks, yet the knowledge of a pre-trained model is tightly coupled with monolithic architecture, which restricts flexible reuse across models of varying scales.

By Jianlu Shen, Fu Feng, Yucheng Xie, Jiaqi Lv, Xin Geng