arXiv Machine Learning By Yangang Ren, Yujie Yan, Zirui Li, Jiaming Guo, Di Zeng, Ji Tao, Lan Yu, Xuesong Tian, Chen Lv

Learning from Mixed-Quality Deployment Experience for Robot Manipulation

Read the original on arXiv Machine Learning →

The paper introduces Predictive Action Chunk Learning (PACL), a method for improving robot manipulation policies using mixed-quality deployment experience. PACL first trains a predictive chunk-level critic to evaluate temporally extended action sequences, then uses the critic’s quality estimates to guide a diffusion actor that learns from both successful and failed rollouts. Experiments on simulated and real robots demonstrate that PACL consistently enhances pretrained policies and outperforms strong imitation learning and offline reinforcement learning baselines.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Jul 7

HALO-WA: Hybrid-Attention Latent-Guided Online Reinforcement Learning for World-Action Models

arXiv:2607. 04265v1 Announce Type: cross Abstract: World-action (WA) models can generate long-horizon action chunks for general-purpose robotic manipulation, but they remain vulnerable to calibration, perception, and contact-dynamics errors in real-world precision tasks, often failing in the final few millimeters of alignment or insertion.

By Angen Ye, Weijie Ke, Xiaofeng Wang, Xinze Chen, Chaojun Ni, Guosheng Zhao, Boyuan Wang, Zheng Zhu, Junjie Xie, Dapeng Zhang
arXiv Machine Learning
Sep 3

Act More, Decide Less: Skill-Guided Adaptive Action Chunking for Long-Horizon LLM Agents

The paper introduces SPACE, a method for enabling large language model agents to emit variable-length action chunks in long-horizon tasks. By distilling chunk-boundary supervision from programmatic skills derived from successful trajectories, SPACE overcomes the tendency of agents to either act one step at a time or commit to overly long sequences. Experiments on ALFWorld and ScienceWorld demonstrate that SPACE raises success rates by 7.0%–31.3% and cuts LLM decision rounds by up to 78.9%.

By Yanting Yang, Can Jin, Jinman Zhao, Jiahao Wu, Yang Zhou, Zhepeng Wang, Zhendong Wang, Mu Zhou, Dimitris N. Metaxas
arXiv Machine Learning
Sep 21

Rollout Total Correlation for Deep Reinforcement Learning

The paper proposes a method for learning task-relevant representations in deep reinforcement learning by maximizing rollout total correlation, which captures the correlation among all learned representations and actions across entire trajectories. It introduces two complementary lower bounds—one generative and one discriminative—along with chunk‑wise mini‑batching to improve this objective, and also proposes an intrinsic reward derived from the learned representation to enhance exploration. Experiments on challenging image‑based simulated control tasks demonstrate improved sample efficiency and robustness to white noise and natural video backgrounds compared to leading baselines.

By Bang You, Huaping Liu, Jan Peters, Oleg Arenz