Source-Lifted Flow Matching for Intervenable Multimodal Imitation
arXiv:2607. 10206v1 Announce Type: cross Abstract: Flow-matching policies are promising for imitation learning because they model complex multimodal action distributions.
The paper investigates how generative behavioral-cloning policies handle multimodal expert behavior, identifying bottlenecks in both latent-variable and action-space policy designs. For latent-variable policies, preserving demonstrated modes depends on action-conditioned latent representations, and excessive posterior-prior regularization can suppress this information. In action-space generative policies, multimodality is limited by the smoothness of the base-to-action transport, requiring either sharp transitions or off-support bridge regions to capture many well-separated modes. Experiments on synthetic navigation and a physical-robot manipulation task confirm these findings, while standard robotic simulation benchmarks show limited conditional multimodality, making deterministic regression competitive.
arXiv:2607. 10206v1 Announce Type: cross Abstract: Flow-matching policies are promising for imitation learning because they model complex multimodal action distributions.
arXiv:2603. 22876v2 Announce Type: replace-cross Abstract: Learning a generalist control policy for robotic manipulation typically relies on large-scale datasets.
arXiv:2609.09210v1 Announce Type: cross Abstract: Teleoperated demonstrations are often multimodal even when the underlying dynamics are nearly deterministic given the executed action. We argue that...
arXiv:2606. 08657v1 Announce Type: cross Abstract: Diffusion-based visuomotor policies operating directly in raw action spaces conflate scene comprehension with trajectory generation within a single denoising process.
arXiv:2505. 04999v2 Announce Type: replace-cross Abstract: Learning robot control policies from demonstrations typically requires action-labeled expert data, which is expensive to collect through teleoperation.
arXiv:2608. 19613v1 Announce Type: cross Abstract: Latent Action Models (LAMs) have emerged as a promising paradigm for enabling robot learning to leverage large-scale unlabeled videos through latent actions that serve as compact surrogates for physical actions.
arXiv:2607. 17257v1 Announce Type: cross Abstract: Diffusion policies have shown strong potential for robotic imitation learning, and recent extensions incorporate additional modalities to improve manipulation performance.
arXiv:2606. 14981v1 Announce Type: cross Abstract: Inference-time steering adapts pre-trained generative robot policies during deployment by verifying candidate actions before execution.
arXiv:2604. 21391v2 Announce Type: replace-cross Abstract: Bridging high-level semantic understanding with low-level physical control remains a persistent challenge in embodied intelligence, stemming from the fundamental spatiotemporal scale mismatch between cognition and action.
arXiv:2602. 09580v4 Announce Type: replace-cross Abstract: Real-world fine-tuning of dexterous manipulation policies remains challenging due to limited real-world interaction budgets and highly multimodal action distributions.
arXiv:2607. 03964v1 Announce Type: cross Abstract: World models are rapidly becoming a core infrastructure for embodied intelligence and interactive agents: they provide controllable simulators in which agents can perceive, act, forecast, and acquire scalable experience.
arXiv:2609.10506v1 Announce Type: cross Abstract: Action-conditioned latent world models predict future visual representations, enabling zero-shot goal-conditioned robot planning and control. However...