arXiv AI By Ayoub Kirouane, Georgios Giaples, Christos Petrocheilos

Measuring Language Transfer in Robot Policies: Adding Greek to a Cosmos3 Vision-Language-Action Policy

Read the original on arXiv AI →

The paper investigates adding Greek language support to a Cosmos3 vision‑language‑action robot policy using only machine‑rephrased instructions and no architectural changes. It finds that many evaluation metrics can give misleading results, and that multilingual training with Greek yields a modest 6.7‑7.1 point advantage over a control, reaching about 40% of English performance. The study also shows that overfitting to a single translator’s phrasing can be mitigated by training on multiple phrasings, while warm‑starting from a language‑adapted world model or unfreezing the text tower actually harms performance.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Sep 22

An Unexpected Robot Policy: Early Evaluations of GPT-6 Astra on RoboDojo and Beyond

arXiv:2609.24170v1 Announce Type: new Abstract: Embodied AI systems are often organized into System 1 and System 2. System 1 is typically a pretrained policy that generates actions at high frequency,...

By Wenbo Zhang, Kaixuan Wang, Yutao Ouyang, Xiaoyu Huang, Liyang Li, Kailun Su, Weiyang Jin, Wenhao Chai, Haotian Liang, Zhiyang Dou, Yue Chen, Tianxing Chen
arXiv Machine Learning
Jun 11

Learning What to Say to Your VLA: Mostly Harmless Vision Language Action Model Steering

arXiv:2606. 12299v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models provide a natural language interface to robot control, but the mapping from language to behavior is often brittle and unintuitive: semantically similar instructions can induce drastically different behaviors, while some capabilities may not be elicitable through prompting alone.

By Hyun Joe Jeong, Gokul Swamy, Andrea Bajcsy