arXiv AI

Stabilizing the Q-Gradient Field for Policy Smoothness in Actor-Critic Methods

arXiv:2601. 22970v2 Announce Type: replace-cross Abstract: Policies learned via continuous actor-critic methods often exhibit erratic, high-frequency oscillations, making them unsuitable for physical deployment.

arXiv AI
Jun 6

Retry Policy Gradients in Continuous Action Spaces

arXiv:2606. 05888v1 Announce Type: new Abstract: Retry-based objectives such as pass@K and max@K optimize the best return obtained from multiple sampled trajectories, and recent work has shown that they can promote exploration without explicit exploration bonuses.

By Soichiro Nishimori, Paavo Parmas
arXiv AI
3d ago

DAMPER: Return-Prioritized Gradient Control for Smooth Policies

DAMPER is a new technique for actor‑critic methods that reduces action oscillation in continuous control tasks. It combines the native actor gradient with a temporal‑consistency gradient using conflict‑conditioned projection and adaptive magnitude control, ensuring the auxiliary component aligns positively with the actor gradient. Experiments on TD3 and SAC across six tasks show that DAMPER consistently lowers oscillation compared to baseline agents and outperforms other methods in most task‑backbone pairs.

By Seokmin Ko, Taewon Goo, Kihyuk Hong