Understanding Multimodality in Generative Behavioral Cloning
Read the original on arXiv AI →The paper investigates how generative behavioral-cloning policies handle multimodal expert behavior, identifying bottlenecks in both latent-variable and action-space policy designs. For latent-variable policies, preserving demonstrated modes depends on action-conditioned latent representations, and excessive posterior-prior regularization can suppress this information. In action-space generative policies, multimodality is limited by the smoothness of the base-to-action transport, requiring either sharp transitions or off-support bridge regions to capture many well-separated modes. Experiments on synthetic navigation and a physical-robot manipulation task confirm these findings, while standard robotic simulation benchmarks show limited conditional multimodality, making deterministic regression competitive.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.