arXiv:2609.25558v1 Announce Type: cross
Abstract: Vision-language-action policies benefit from geometric supervision, but current-frame geometry alone does not explicitly describe the changes associa...
By Jinu Pahk, Jesoon Kang, Taegeon Park, Jisu An, Soo Min Kimm, Jaejoon Kim, Byoung-Tak Zhang
arXiv:2608. 11739v1 Announce Type: cross Abstract: The prevailing recipe for Vision-Language-Action (VLA) models couples a pretrained VLM with a separately trained flow-matching action expert.
By Yicheng Liu, Zibin Dong, Baijun Ye, Tianyuan Yuan, Tao Jiang, Anqi Yang, Shicheng Cao, Haonan Liu, Yue Sun, Zihan Guo, Xiao Liu, Dong Ke, Changxun Pan, Chenru Wu, Tailai Cheng, Xiaoshu Ren, Xinlei Zhang, Jianning Cui, Zijie Zhao, Haoyu Zhang, Kaiming Xu, Haodong Yang, Bowen Zhang, Jiahui Niu, Shaoting Zhu, Shiduo Zhang, Hang Zhao
arXiv:2608.29208v1 Announce Type: cross
Abstract: Vision-Language-Action (VLA) models, built upon Vision-Language Models (VLMs), have significantly enhanced robotic capabilities by leveraging interne...
By Sunghwan Han, Youngtae Han, Youngmin Yi
Task-Prototype Guided Flow Matching (TP-Flow) is a few‑shot manipulation framework that transforms support demonstrations into structured task‑prototype tokens to guide both the initial flow prior and the velocity field. It uses symmetric cross‑attention with learnable queries to extract phase‑level prototypes, parameterizes a task‑adaptive initial distribution, and injects prototype information through gated adaptive normalization. TP‑Flow is trained with an episodic support‑query objective and prototype contrastive regularization, achieving high success rates on the LEROBOT‑ARM‑SO101 platform while maintaining real‑time execution and low latency.
By Yizhao Wang, Guantao Zhang, Jingbo Wang
arXiv:2603. 07523v3 Announce Type: replace Abstract: Transferring knowledge by fine-tuning large-scale pre-trained networks has become a standard paradigm for downstream tasks, yet the knowledge of a pre-trained model is tightly coupled with monolithic architecture, which restricts flexible reuse across models of varying scales.
By Jianlu Shen, Fu Feng, Yucheng Xie, Jiaqi Lv, Xin Geng
arXiv:2609.37250v1 Announce Type: cross
Abstract: World-action models (WAMs) couple future visual-state prediction with action generation. By adapting video generators or image-editing models pretrai...
By Yang Zhang, Jiangyuan Zhao, Chenyou Fan, Jiayu Hu, Xiu Yuan, Chenjia Bai, Xiu Li
Frozen perception foundation models encode rich geometric, semantic, and dynamic knowledge. Yet narrow conditioning interfaces may attenuate task-relevant cues, while static fusion cannot adjust expert contributions to each scene.
The paper proposes a method to train efficient multi‑task manipulation policies by distilling knowledge from single‑task Conditional Flow Matching (CFM) experts. Instead of training separate models for each task, the authors transfer the experts’ learned velocity fields into a shared policy, combining this distillation signal with the original CFM objective. Experiments on RLBench demonstrate that this approach improves multi‑task performance while keeping the model size fixed, avoiding the need for larger capacity or performance drops seen with naive concatenated training.
By Shreya Deshmukh, Imen Mahdi, Nick Heppert, Abhinav Valada
The paper introduces a framework for Flow‑Matching Vision‑Language‑Action (VLA) models that allows independent adjustment of backbone depth, action expert depth, and denoising steps. Lightweight Exit Transformers are added at intermediate layers to enable early exits, and a KV Cache synthesis mechanism manages skipped layers so the action expert can exit deeper than the backbone. Experiments on SmolVLA and π0.5 across LIBERO and Meta‑World show that joint tuning of these compute axes reduces latency by 79.2 % and FLOPs by 31.8 %, while improving mean success rate by 5.6 %.
By Riccardo Andrea Izzo, Rimvydas Rubavicius, Gianluca Bardaro, Subramanian Ramamoorthy, Matteo Matteucci, Alessandro Suglia
The paper presents a method for cross‑architecture knowledge distillation from a fine‑tuned DINOv2 Vision Transformer teacher to a lightweight bidirectional Visual State Space Model (LVSSM) student for tea leaf disease classification. By addressing training‑stability issues with a progressive convolutional stem and gated selective‑scan block, the 4.45 M‑parameter student achieves a mean test accuracy of 95.41%—a 3.09‑point improvement over the teacher’s 92.32%—while using only one‑fifth of the teacher’s parameters. Ablation studies show that simple logit‑level distillation outperforms intermediate feature alignment, and the gains are specific to students that start below the teacher’s performance.
By Zibo Zhou, Zongsen Qiu, Rui Chen, Yujie Yao, Yue Zhou, Jianjun Wang
Dynin‑Robotics introduces an omnimodal masked‑diffusion backbone, Dynin‑Omni, that jointly represents language, visual observations, goals, and actions as discrete tokens. By conditioning on different spans, the same model learns action prediction, next‑observation prediction, goal‑state prediction, and trajectory‑to‑instruction reconstruction, enabling test‑time scaling through goal prediction and action‑candidate evaluation. The system, pretrained on 1.33 million trajectories from 48 Open X‑Embodiment datasets, achieves competitive performance on LIBERO, zero‑shot LIBERO‑Plus, and a 78.4 % success rate on a Franka Research 3 robot, while a block‑parallel implementation speeds up action decoding by up to 29.2×.
By Hoeun Lee, Jaeik Kim, Jusang Oh, Jinhyeok Kim, Geon Choi, Hyeonggeun Kim, Jaeyoung Do
arXiv:2609.10321v1 Announce Type: new
Abstract: Knowledge distillation offers an efficient route to transfer a task-adapted vision-language teacher to a compact student. The training target in curren...
By Hongyuan Zhang, Xianda Guo, Yanlun Peng, Qianlong Yang, Yubin Guo, Pinhan Fu, Mulin Chen, Xiaozhen Qiao, Ping Luo