arXiv Machine Learning

SeFA-Policy: Fast and Accurate Visuomotor Policy Learning with Selective Flow Alignment

arXiv:2511. 08583v2 Announce Type: replace-cross Abstract: Developing efficient and accurate visuomotor policies poses a central challenge in robotic imitation learning.

arXiv Machine Learning
Aug 31

FlowCorrect: Efficient Interactive Correction of Generative Flow Policies for Robotic Manipulation

FlowCorrect is a modular interactive imitation learning method that allows real‑time adaptation of generative flow‑matching manipulation policies using sparse, relative human corrections. During task execution, a human provides brief corrective pose nudges through a lightweight VR interface, and FlowCorrect locally adapts the policy without retraining the backbone, maintaining performance on previously learned scenarios. Experiments on a real‑world robot across four tabletop tasks show that, with a low correction budget, FlowCorrect achieves an 80% success rate on previously failed cases while preserving performance on solved scenarios.

By Edgar Welte, Yitian Shi, Rosa Wolf, Maximillian Gilles, Rania Rayyes
arXiv AI
Jul 7

High-Fidelity One-Step Generative Visuomotor Policy via Recursive Correction, Frequency Consistency, and Contrastive Flow Matching

arXiv:2607. 03865v1 Announce Type: cross Abstract: Generative models such as diffusion and flow matching have advanced robotic visuomotor policies by modeling multimodal action distributions, but their multi-step sampling or ODE solving introduces inference latency.

By Yuran Chen, Xinye Cai, Zhonglin Gong, Yang Huang
arXiv Machine Learning
Jun 8

SERNF: Sample-Efficient Real-World Dexterous Policy Fine-Tuning via Action-Chunked Critics and Normalizing Flows

arXiv:2602. 09580v4 Announce Type: replace-cross Abstract: Real-world fine-tuning of dexterous manipulation policies remains challenging due to limited real-world interaction budgets and highly multimodal action distributions.

By Chenyu Yang, Denis Tarasov, Davide Liconti, Romain Guntz, Hehui Zheng, Robert K. Katzschmann
arXiv AI
2d ago

DriftOPD: Sequence-Level Reverse-KL Distillation for One-Step VLA Policies

DriftOPD is a teacher‑free, rollout‑free framework that performs sequence‑level on‑policy distillation of continuous Vision‑Language‑Action (VLA) action experts. It decomposes the sequence‑level reverse‑KL divergence into a chunk‑level reverse‑KL term and a future‑potential term, optimizing them with a one‑step drifting objective and a Q‑function critic learned from offline demonstrations. Experiments on multiple VLA architectures in simulation and real‑world manipulation show that DriftOPD outperforms existing one‑step distillation baselines while matching the task success of multi‑step teacher policies.

By Youngjun Jun, Kyumin Choi, Youngmin Kim, Seonghyun Jin, Sunwoo Park, Jangho Park, Jong Chul Ye
arXiv AI
Sep 24

Distillation for Efficient Multitask Manipulation Policies via Conditional Flow Matching

The paper proposes a method to train efficient multi‑task manipulation policies by distilling knowledge from single‑task Conditional Flow Matching (CFM) experts. Instead of training separate models for each task, the authors transfer the experts’ learned velocity fields into a shared policy, combining this distillation signal with the original CFM objective. Experiments on RLBench demonstrate that this approach improves multi‑task performance while keeping the model size fixed, avoiding the need for larger capacity or performance drops seen with naive concatenated training.

By Shreya Deshmukh, Imen Mahdi, Nick Heppert, Abhinav Valada