arXiv AI By Markus Knauer, Edoardo Fiorini, Maximilian M\"uhlbauer, Stefan Schneyer, Promwat Angsuratanawech, Florian Samuel Lay, Timo Bachmann, Samuel Bustamante, Korbinian Nottensteiner, Freek Stulp, Alin Albu-Sch\"affer, Jo\~ao Silv\'erio, Thomas Eiband

MOMO: A framework for seamless physical, verbal, and graphical robot skill learning and adaptation

Read the original on arXiv AI →

The paper introduces MOMO, a framework that allows industrial robots to be adapted by non-experts using kinesthetic touch, natural language, and a graphical web interface. It combines energy‑based intention detection, a tool‑based LLM for safe language adaptation, Kernelized Movement Primitives for motion encoding, probabilistic Virtual Fixtures for guided demonstrations, and ergodic control for surface finishing. The authors validate the system on a 7‑DoF torque‑controlled robot at the Automatica 2025 trade fair, showing its practical applicability in industrial settings.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 9

GEAR-VLA: Learning Geometry-Aware Action Representations for Generalizable Robotic Manipulation

arXiv:2606. 08530v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models achieve strong benchmark performance but still struggle in real-world deployment with unseen objects, background shifts, and different robot embodiments.

By Yuan Zhang, Shiqi Zhang, Yedong Shen, Shuai Dong, Jiajun Deng, Xin Zhang, Yuxuan Gao, Jiajia Wu, Xin Nie, Zhiyuan Cheng, Jianmin Ji, Yanyong Zhang, Xingyi Zhang, Jia Pan
arXiv AI
Jun 16

Action with Visual Primitives

arXiv:2605. 22183v3 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models have emerged as a promising paradigm for generalist robotic manipulation.

By Weilong Guo, Yuchen Wang, Renping Zhou, Yunfeng Zhang, Rui Fang, Yuyang Pang, Wenda Xu, Gao Huang