arXiv AI By Lorenzo Mazza, Massimiliano Datres, Ariel Rodriguez, Sebastian Bodenstedt, Gitta Kutyniok, Stefanie Speidel

Understanding Multimodality in Generative Behavioral Cloning

Read the original on arXiv AI →

The paper investigates how generative behavioral-cloning policies handle multimodal expert behavior, identifying bottlenecks in both latent-variable and action-space policy designs. For latent-variable policies, preserving demonstrated modes depends on action-conditioned latent representations, and excessive posterior-prior regularization can suppress this information. In action-space generative policies, multimodality is limited by the smoothness of the base-to-action transport, requiring either sharp transitions or off-support bridge regions to capture many well-separated modes. Experiments on synthetic navigation and a physical-robot manipulation task confirm these findings, while standard robotic simulation benchmarks show limited conditional multimodality, making deterministic regression competitive.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Aug 21

What Matters for Latent Actions in Robot Learning

arXiv:2608. 19613v1 Announce Type: cross Abstract: Latent Action Models (LAMs) have emerged as a promising paradigm for enabling robot learning to leverage large-scale unlabeled videos through latent actions that serve as compact surrogates for physical actions.

By Xizhou Bu, Qingda Hu, Lei Zhou, Lingfeng Zhang, Yingbo Tang, Zihao Liu, Xinyi Tao, Zhiqiang Ma, Qingqiu Huang, Chufeng Tang, Hongbo Wang, Jing Zhang, Jiayi Ma, Hangjun Ye, Wei Li, Xiaoshuai Hao