arXiv AI By Seokmin Ko, Taewon Goo, Kihyuk Hong

DAMPER: Return-Prioritized Gradient Control for Smooth Policies

Read the original on arXiv AI →

DAMPER is a new technique for actor‑critic methods that reduces action oscillation in continuous control tasks. It combines the native actor gradient with a temporal‑consistency gradient using conflict‑conditioned projection and adaptive magnitude control, ensuring the auxiliary component aligns positively with the actor gradient. Experiments on TD3 and SAC across six tasks show that DAMPER consistently lowers oscillation compared to baseline agents and outperforms other methods in most task‑backbone pairs.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Jun 11

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies

arXiv:2605. 03065v2 Announce Type: replace Abstract: Generative control policies (GCPs), such as diffusion- and flow-based control policies, have emerged as effective parameterizations for robot learning.

By Sarvesh Patil, Mitsuhiko Nakamoto, Manan Agarwal, Shashwat Saxena, Jesse Zhang, Giri Anantharaman, Cleah Winston, Chaoyi Pan, Douglas Chen, Nai-Chieh Huang, Zeynep Temel, Oliver Kroemer, Sergey Levine, Abhishek Gupta, Hongkai Dai, Paarth Shah, Max Simchowitz