Vision-Language-Action (VLA) models have demonstrated strong capabilities in robotic manipulation by integrating visual perception, language understanding, and robot action generation. Existing research has primarily focused on improving model architectures, training strategies, and dataset scale, while little attention has been paid to how demonstrations are collected and organized.
Vision-Language-Action (VLA) models are commonly treated as end-to-end action policies conditioned on natural-language task descriptions. In practice, however, their behavior often depends sharply on how the instruction is phrased, suggesting that language is not merely a task label but an optimizable conditioning input.
Flow matching is a powerful tool for generative modeling, but emerging applications in robotics, planning, and physics require inference-time constraints on generated outputs. Such constraints are often complex and highly nonlinear.
arXiv:2607. 02426v1 Announce Type: cross Abstract: Federated learning (FL) enables collaborative model training across distributed devices without sharing raw data, making it suitable for privacy-sensitive robotic sensing applications.
By Quoc Bao Phan, Tuy Tan Nguyen
arXiv:2603. 06001v2 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models enable robots to perform manipulation tasks directly from natural language instructions and are increasingly viewed as a foundation for generalist robotic policies.
By Ninghao Zhang, Bin Zhu, Shijie Zhou, Jingjing Chen
arXiv:2604. 26689v4 Announce Type: replace-cross Abstract: Compositional machine-learning (ML) systems assemble runtime behavior from libraries of independently re-trained capability modules.
By Xue Qin, Simin Luan, Cong Yang, Zhijun Li
arXiv:2607. 02417v1 Announce Type: cross Abstract: Autonomous robots often need to move their camera before they can act: to inspect an object, reveal an occluded region, or obtain a view that responds to a user's intent.
By Boyang Sun, Jiajie Li, Yung-Hsu Yang, Chenyangguang Zhang, Tim Engelbracht, Sunghwan Hong, Cesar Cadena, Marc Pollefeys, Hermann Blum
arXiv:2604. 03497v2 Announce Type: replace-cross Abstract: Vision-language-model (VLM)-guided reinforcement learning (RL) has recently attracted significant attention for it, replacing brittle hand-crafted rewards with semantically grounded signals; however, deploying such simulation-trained policies on real vehicles remains a fundamental challenge, because they rely on simulator-native observations and simulator-coupled action semantics with no counterpart on physical hardware.
By Zilin Huang, Zhengyang Wan, Zihao Sheng, Boyue Wang, Junwei You, Sikai Chen
arXiv:2607. 01824v1 Announce Type: new Abstract: Crowdsourced fact-checking systems have been adopted by major social media companies such as X, Meta, TikTok and Google with the aim of combating misleading information at scale without relying on centralized editorial control.
By Nikil Roashan Selvam, Jay Baxter, Sophie Hilgard, Brad Miller, Keith Coleman, Ellen Vitercik, Sanmi Koyejo
arXiv:2607. 01410v1 Announce Type: cross Abstract: Sim2real transfer for robot policy learning suffers due to mismatch between simulation and reality.
By Yunfu Deng, Josiah P. Hanna
arXiv:2607. 02222v1 Announce Type: cross Abstract: Vision-Language Navigation has increasingly emphasized high-level instruction reasoning, memory, global map construction, and instruction decomposition, while the low-level action representation remains comparatively underexplored.
By Haokun Liu, Zhaoqi Ma, Yicheng Chen, Wentao Zhang, Masaki Kitagawa, Zicen Xiong, Jinjie Li, Moju Zhao
arXiv:2503. 24009v3 Announce Type: replace-cross Abstract: Realistic simulation is critical for applications ranging from robotics to animation.
By Mikel Zhobro, Andreas Ren\'e Geist, Georg Martius
arXiv:2607. 01766v1 Announce Type: new Abstract: LLM agents are increasingly used to translate natural language into 3D scenes in a procedural way, but existing systems focus on static output.
By Chunjiang Liu, Xiaoyuan Wang, Haoyu Chen, Yizhou Zhao, Ming-Hsuan Yang, L\'aszl\'o A. Jeni
arXiv:2605. 00412v3 Announce Type: replace Abstract: World models have recently re-emerged as a central paradigm for embodied intelligence, robotics, autonomous driving, and model-based reinforcement learning.
By Sen Cui, Jingheng Ma
arXiv:2510. 04391v5 Announce Type: replace Abstract: Mental imagery vividness is a stable individual trait, yet whether imagined scenarios share relational structure across human and synthetic large language model (LLM) populations remains unknown.
By Saurabh Ranjan, Brian Odegaard
arXiv:2607. 01754v1 Announce Type: new Abstract: On-policy exploration is a crucial component for training robust Vision-Language Navigation agents, as it exposes the policy to a broader state distribution.
By Sung June Kim, Sangpil Kim, Honglak Lee
arXiv:2510. 06288v4 Announce Type: replace Abstract: Today's AI models learn primarily through mimicry and refining, so it is not surprising that they struggle to solve problems beyond the limits set by existing data.
By Raj Ghugare, Roger Creus Castanyer, Catherine Ji, Kathryn Wantlin, Jin Schofield, Karthik Narasimhan, Benjamin Eysenbach
arXiv:2607. 01938v1 Announce Type: cross Abstract: Manipulating fast and dynamically moving targets in unstructured 3D environments remains challenging for embodied AI.
By Peng Yun, Shouwang Huang, Hao Li, Jinxi Li, Jianan Wang, Bo Yang
arXiv:2607. 01435v1 Announce Type: cross Abstract: Significant advancement of immersive technologies such as Virtual and Augmented Reality (VR/AR) and their integration into diverse aspects of modern life need authentication interfaces that are secure, intuitive, and compatible with embodied interaction.
By Neda Abdolrahimi, Thiru Siddharth, Frank Sicongchen, Vir V Phoha
arXiv:2607. 02432v1 Announce Type: new Abstract: Scalable and reliable grading of command-line examinations remains a challenge in computing education, where rising enrolments make manual marking difficult and rule-based autograders cannot handle partial credit, equivalent solutions, or syntactic variation.
By Manuel Alonso-Carracedo, Ruben Fernandez-Boullon, Pedro Celard, Francisco J. Rodriguez-Martinez, Lorena Otero-Cerdeira