arXiv Machine Learning By Saki Omi, Hyo-Sang Shin, Namhoon Cho, Antonios Tsourdos, Miguel A. Olivares-Mendez

Robust Recurrent Reinforcement Learning under Evolving Hidden Disturbances with Application to Rover Wheel Slip

Read the original on arXiv Machine Learning →

The paper studies recurrent Twin Delayed Deep Deterministic Policy Gradient (TD3) agents in environments with evolving hidden disturbances, focusing on how observation history, action history, history length, and network structure influence performance. Three recurrent architectures are compared under controlled disturbances, revealing that action history is crucial when responses depend on prior actions and that a unified temporal sequence of action-observation pairs outperforms separate branches. The authors introduce H‑TD3, which reuses actor-generated recurrent states to initialize the critic, and demonstrate that these architectures excel in a rover wheel‑slip simulation, with policies trained on temporally structured disturbances transferring better to unseen slip dynamics.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
2d ago

DriftOPD: Sequence-Level Reverse-KL Distillation for One-Step VLA Policies

DriftOPD is a teacher‑free, rollout‑free framework that performs sequence‑level on‑policy distillation of continuous Vision‑Language‑Action (VLA) action experts. It decomposes the sequence‑level reverse‑KL divergence into a chunk‑level reverse‑KL term and a future‑potential term, optimizing them with a one‑step drifting objective and a Q‑function critic learned from offline demonstrations. Experiments on multiple VLA architectures in simulation and real‑world manipulation show that DriftOPD outperforms existing one‑step distillation baselines while matching the task success of multi‑step teacher policies.

By Youngjun Jun, Kyumin Choi, Youngmin Kim, Seonghyun Jin, Sunwoo Park, Jangho Park, Jong Chul Ye
arXiv Computer Vision
Sep 18

Learning Foresight without Explicit Trajectories for 3D Diffusion Policies

The paper introduces Movement Trend Guidance, a method that equips 3D diffusion policies with foresight by learning a compact latent representation of interaction evolution from a brief observation history. This latent, supervised by sparse future gripper states during training, serves as future-oriented conditioning during inference, enhancing action generation without adding explicit planning. The approach improves performance on RoboTwin2.0, LIBERO-40, and DexArt benchmarks, achieving higher success rates across multiple tasks.

By Zhongbo Zhang, Zaibin Zhang, Yifan Wang, Changbo Yan, Lijun Wang, Huchuan Lu