arXiv Machine Learning By Adrian Ly, Richard Dazeley, Peter Vamplew, Sunil Aryal, Francisco Cruz

Memory Merge DQN: Sensitivity Weighted Target Updates for Stable Value Learning

Read the original on arXiv Machine Learning →

arXiv:2607. 19397v1 Announce Type: new Abstract: Deep Q-networks use target networks to stabilise bootstrapped value learning, but the standard hard copy update also introduces a tradeoff.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 25

Aftab: A Progressive Design Study of Visual Encoders and Value Estimation for Replay-Free Parallelized Q-Learning

The paper presents Aftab, a new architecture for replay‑free parallelized Q‑learning that systematically explores visual encoder designs, multiplicative feature interactions, and value‑estimation strategies. Through a three‑phase study on Atari‑57, the authors compare eight convolutional encoders, integrate Hadamax‑style interactions, and evaluate categorical‑dueling, ensemble‑dueling, and combined configurations, ultimately achieving a higher human‑normalized score than the baseline PQN. Aftab is also evaluated on Procgen Hard, showing improved terminal IQM and learning‑curve area, and the full framework is released as open source.

By Taha Shieenavaz, Shabnam Zareshahraki, Loris Nanni
arXiv AI
6d ago

A Progressive Design Study of Visual Encoders and Value Estimation for Replay-Free Parallelized Q-Learning

The paper presents a three‑phase study of visual encoders and value‑estimation methods for replay‑free parallelized Q‑learning within the Parallelized Q‑Network framework. It compares eight convolutional encoder topologies, adds Hadamax‑style multiplicative interactions and pooling, and evaluates categorical‑dueling, ensemble‑dueling, and categorical ensemble‑dueling value‑estimation configurations. The resulting architecture, Aftab, outperforms a baseline PQN on Atari‑57 and shows improved performance on Procgen Hard, demonstrating that visual topology, multiplicative representation, and downstream value‑estimation design significantly influence replay‑free Q‑learning when considered alongside computational complexity.

By Taha Shieenavaz, Shabnam Zareshahraki, Loris Nanni