The paper presents Aftab, a new architecture for replay‑free parallelized Q‑learning that systematically explores visual encoder designs, multiplicative feature interactions, and value‑estimation strategies. Through a three‑phase study on Atari‑57, the authors compare eight convolutional encoders, integrate Hadamax‑style interactions, and evaluate categorical‑dueling, ensemble‑dueling, and combined configurations, ultimately achieving a higher human‑normalized score than the baseline PQN. Aftab is also evaluated on Procgen Hard, showing improved terminal IQM and learning‑curve area, and the full framework is released as open source.
By Taha Shieenavaz, Shabnam Zareshahraki, Loris Nanni
arXiv:2608. 07335v1 Announce Type: cross Abstract: Recent advancements in deep reinforcement learning have increasingly favored simplified, highly parallelized paradigms.
By Taha Shieenavaz, Shabnam Zareshahraki, Loris Nanni
arXiv:2607. 19397v1 Announce Type: new Abstract: Deep Q-networks use target networks to stabilise bootstrapped value learning, but the standard hard copy update also introduces a tradeoff.
By Adrian Ly, Richard Dazeley, Peter Vamplew, Sunil Aryal, Francisco Cruz
arXiv:2410.14606v3 Announce Type: replace
Abstract: Learning from a stream of experience as it arrives, also known as streaming learning, is a core part of natural learning. However, reliable streami...
By Mohamed Elsayed, Elena Sorina Lupu, Gautham Vasan, A. Rupam Mahmood
arXiv:2608. 02034v1 Announce Type: new Abstract: Multi-step returns accelerate reward propagation in off-policy reinforcement learning, but couple the evaluation of each decision to the suboptimal logged actions that follow it, inducing a pessimistic bias that grows with the horizon.
By Abdelghani Ghanem, Mounir Ghogho
arXiv:2609.06421v1 Announce Type: cross
Abstract: Batch normalization (BN) substantially improves sample efficiency in continuous-control actor-critic methods such as CrossQ, yet recent studies repor...
By Daniel Palenicek, Mikael Henaff, Scott Fujimoto, Koustuv Sinha