Hugging Face Trending Papers

AIDE: Automated Instruction via Distilled Expertise for Reference-Free Motor Skill Coaching

Generating natural-language coaching feedback on motor skills can accelerate learning, yet expert coaches are scarce and expensive. Existing reference-based methods require expert demonstrations at both training and inference time, limiting practical deployment.

arXiv AI
Sep 24

Distillation for Efficient Multitask Manipulation Policies via Conditional Flow Matching

The paper proposes a method to train efficient multi‑task manipulation policies by distilling knowledge from single‑task Conditional Flow Matching (CFM) experts. Instead of training separate models for each task, the authors transfer the experts’ learned velocity fields into a shared policy, combining this distillation signal with the original CFM objective. Experiments on RLBench demonstrate that this approach improves multi‑task performance while keeping the model size fixed, avoiding the need for larger capacity or performance drops seen with naive concatenated training.

By Shreya Deshmukh, Imen Mahdi, Nick Heppert, Abhinav Valada
arXiv Computer Vision
Sep 21

DexPIE: Stable Dexterous Policy Improvement from Real-World Experience

DexPIE is a post‑training framework that improves dexterous manipulation policies using real‑world experience. It introduces a dexterous‑hand‑adapted intervention system and multi‑stage DAgger‑style data collection to enhance exploration, aligns training and inference to reduce distribution shift, and conditions the policy on a continuous optimality indicator for fine‑grained data quality use. In three real‑world tasks, DexPIE boosts success rates by 37.3% over a demonstration‑based baseline, outperforming all other methods and showing stronger robustness.

By Ruizhe Liao, Wenrui Chen, Liangji Zeng, Haoran Lin, Fan Yang, Kailun Yang, Yaonan Wang
arXiv AI
Aug 26

Disentangled Skill Representations for Predictive Human Modeling

The paper introduces Skill Abstraction with Interpretable Latents (SAIL), a method that models human skill as a persistent, multi‑dimensional construct inferred from naturalistic behavior over time. SAIL produces a robust skill embedding that blends expert and novice bases, learns transferable subskills through counterfactual subskill swaps, and supports skill‑informed behavior prediction across various in‑domain contexts. Experiments on racing and baseball demonstrate that SAIL achieves strong predictive performance, improves behaviorally grounded disentanglement compared to baselines, and enhances downstream AI coaching outcomes.

By Mariah Schrum, Deepak Gopinath, Srijan Srivatsa, Guy Rosman, Tiffany Chen
arXiv Computer Vision
6d ago

InternW0-$\Delta$: A World Action Model Bridging Predictive Dynamics and Actions with 20K+ Hours of Open Data

InternW0-Δ is a unified World Action Model that integrates pretrained visual dynamics, scene semantics, 4D geometry, and motion priors within a Mixture-of-Transformers framework to generate robot actions. It leverages a frozen VLM for semantic guidance, a 4D foundation model for geometric priors, and introduces Causal Imprint to learn future-relevant scene changes without future-video rollout. The model is pretrained on a newly curated 20K‑hour heterogeneous corpus of robot and human demonstrations, achieving superior performance on simulation benchmarks and real‑robot platforms.

By Xingyu Miao, Zizun Li, Baole Fang, Kaiwen Song, Tenghui Wang, Hanxue Zhang, Yating Wang, Xudong Li, Yuping He, Xueyuan Wei, Chao Gao, Xijie Yang, Yingxiang Xu, Kerui Ren, Wenqi Guo, Jianjun Zhou, Xinzhe Wang, Weiguang Zhao, Ni Yang, Zetao Cai, Yufei Xue, Hengjie Li, Zeyu He, Yuanzhen Zhou, Rong Fu, Jianyang Zhang, Siwei Cui, Fuxian Huang, Yunsong Zhou, Xing Gao, Yifei Yao, Qiaojun Yu, Kailin Li, Ming Zhou, Mu Huang, Xinyue Li, Wenze Cui, Bingqi Jiang, Xueyue Zhu, Junting Dong, Haoyu Guo, Tao Lu, Mulin Yu, Bowen Zhou, Bin Zhao, Tianfan Xue, Weinan Zhang, Chunhua Shen