arXiv AI

Privileged observations enable rapid and reliable policy discovery directly in the physical world

arXiv Computer Vision
Sep 18

Learning Foresight without Explicit Trajectories for 3D Diffusion Policies

The paper introduces Movement Trend Guidance, a method that equips 3D diffusion policies with foresight by learning a compact latent representation of interaction evolution from a brief observation history. This latent, supervised by sparse future gripper states during training, serves as future-oriented conditioning during inference, enhancing action generation without adding explicit planning. The approach improves performance on RoboTwin2.0, LIBERO-40, and DexArt benchmarks, achieving higher success rates across multiple tasks.

By Zhongbo Zhang, Zaibin Zhang, Yifan Wang, Changbo Yan, Lijun Wang, Huchuan Lu
arXiv AI
Aug 13

TMRL: Diffusion Timestep-Modulated Pretraining Enables Exploration for Efficient Policy Finetuning

arXiv:2605. 12236v2 Announce Type: replace-cross Abstract: Fine-tuning pre-trained robot policies with reinforcement learning (RL) often inherits the bottlenecks introduced by pre-training with behavioral cloning (BC), which produces narrow action distributions that lack the coverage necessary for downstream exploration.

By Matthew M. Hong, Jesse Zhang, Anusha Nagabandi, Abhishek Gupta
arXiv AI
2d ago

Measuring the Stability Assumption Behind Action Chunking

The paper investigates how small action errors evolve when using action chunking in behavioural cloning. By injecting errors at each state and observing their growth under open‑loop (no replanning) and closed‑loop (replanning) regimes, the authors classify states as contracting, expanding, or unresolved. Across twelve manipulation tasks, they find that stable states are rare, error amplification is common, and that short‑horizon fitting can overestimate long‑horizon propagation. Predictors trained on camera and proprioceptive data can recover open‑loop stability but only partially capture closed‑loop dynamics, indicating that standard imitation learning does not reliably produce policies that contract errors when perturbed.

By Aryan Goyal
arXiv AI
Sep 2

Explore More, Drift Less: Outcome-Only Reinforcement Learning Can Suffice for Long-Horizon Interactive Agents

The paper proposes CANOPY, a minimalist reinforcement learning protocol that addresses two common pitfalls—signal starvation and policy drift—in outcome‑only RL for long‑horizon interactive tasks. By scaling same‑task exploration, keeping updates on‑policy, and anchoring updates with KL divergence, CANOPY enables a Qwen3‑14B agent to achieve top leaderboard results on the AppWorld coding benchmark without auxiliary supervision or elaborate scaffolding. The approach also improves performance on SWE‑bench for a Qwen3.5‑9B model.

By Liming Pu, Xiaoxia Li, Yifu Liu, Teng Cao, Bin Yang
arXiv Machine Learning
Jun 11

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies

arXiv:2605. 03065v2 Announce Type: replace Abstract: Generative control policies (GCPs), such as diffusion- and flow-based control policies, have emerged as effective parameterizations for robot learning.

By Sarvesh Patil, Mitsuhiko Nakamoto, Manan Agarwal, Shashwat Saxena, Jesse Zhang, Giri Anantharaman, Cleah Winston, Chaoyi Pan, Douglas Chen, Nai-Chieh Huang, Zeynep Temel, Oliver Kroemer, Sergey Levine, Abhishek Gupta, Hongkai Dai, Paarth Shah, Max Simchowitz
arXiv Machine Learning
Sep 25

Uncertainty-Gated Exploration Noise Suppresses Task Collapse in Online RL Fine-Tuning of a Flow-Matching Vision-Language-Action Policy

The paper investigates task collapse—a failure mode where online RL fine‑tuning of a pretrained flow‑matching vision‑language‑action policy erodes performance on individual tasks—using a 450M‑parameter SmolVLA policy on LIBERO‑10. Three exploration‑noise strategies are compared: a fixed noise scale, a learned noise network, and an uncertainty‑gated controller that reallocates exploration based on novelty and competence signals without task labels. The uncertainty‑gated controller prevents task collapse across all tested seeds, whereas the other two approaches consistently cause collapse, demonstrating its effectiveness in preserving task performance during fine‑tuning.

By Mehmet Turan Yard{\i}mc{\i}, Yunus Emre \c{C}o\u{g}urcu
arXiv Machine Learning
Jun 2

Coherent Off-Policy Improvement of Large Behavior Models with Learned Rewards

arXiv:2606. 02194v1 Announce Type: new Abstract: Distilling expert demonstration data into large generative models using behavioral cloning is a scalable approach to learning capable policies for robotic control, particularly for dexterous manipulation.

By Christian Scherer, Joe Watson, Theo Gruner, Daniel Palenicek, Ingmar Posner, Jan Peters
arXiv AI
4d ago

Neuro-Symbolic Computer Use: Learning Reusable Policies for Reliable and Efficient Execution

The paper introduces neuro‑symbolic computer use, a method that learns reusable policies to execute recurring computer workflows efficiently. Instead of re‑planning each run, the learned policy encodes stable decisions (ordering, variables, loops, branches) into executable code while delegating observation‑dependent decisions to neural models. Using neuro‑symbolic policy iteration, the approach iteratively refines the policy from a single agent trajectory, diagnoses failures, and revises the code with a coding model, achieving superior Pass^3 scores and significant reductions in per‑run cost and latency on OSWorld‑Verified and ScienceBoard benchmarks.

By Hyewon Suh, Thanh Minh Nguyen, Chih-Lun Lee, Darrow Hartman, Lizhao Liu, Xin Eric Wang, Ang Li, Jiachen Yang