The paper presents a three‑phase study of visual encoders and value‑estimation methods for replay‑free parallelized Q‑learning within the Parallelized Q‑Network framework. It compares eight convolutional encoder topologies, adds Hadamax‑style multiplicative interactions and pooling, and evaluates categorical‑dueling, ensemble‑dueling, and categorical ensemble‑dueling value‑estimation configurations. The resulting architecture, Aftab, outperforms a baseline PQN on Atari‑57 and shows improved performance on Procgen Hard, demonstrating that visual topology, multiplicative representation, and downstream value‑estimation design significantly influence replay‑free Q‑learning when considered alongside computational complexity.
By Taha Shieenavaz, Shabnam Zareshahraki, Loris Nanni
arXiv:2608. 07335v1 Announce Type: cross Abstract: Recent advancements in deep reinforcement learning have increasingly favored simplified, highly parallelized paradigms.
By Taha Shieenavaz, Shabnam Zareshahraki, Loris Nanni
arXiv:2607. 19397v1 Announce Type: new Abstract: Deep Q-networks use target networks to stabilise bootstrapped value learning, but the standard hard copy update also introduces a tradeoff.
By Adrian Ly, Richard Dazeley, Peter Vamplew, Sunil Aryal, Francisco Cruz
arXiv:2410.14606v3 Announce Type: replace
Abstract: Learning from a stream of experience as it arrives, also known as streaming learning, is a core part of natural learning. However, reliable streami...
By Mohamed Elsayed, Elena Sorina Lupu, Gautham Vasan, A. Rupam Mahmood
arXiv:2608. 02034v1 Announce Type: new Abstract: Multi-step returns accelerate reward propagation in off-policy reinforcement learning, but couple the evaluation of each decision to the suboptimal logged actions that follow it, inducing a pessimistic bias that grows with the horizon.
By Abdelghani Ghanem, Mounir Ghogho
arXiv:2609.06421v1 Announce Type: cross
Abstract: Batch normalization (BN) substantially improves sample efficiency in continuous-control actor-critic methods such as CrossQ, yet recent studies repor...
By Daniel Palenicek, Mikael Henaff, Scott Fujimoto, Koustuv Sinha
arXiv:2602. 12643v2 Announce Type: replace-cross Abstract: We present Unified Latent Dynamics (ULD), a novel reinforcement learning algorithm that unifies the efficiency of model-free methods with the representational strengths of model-based approaches, without incurring planning overhead.
By Jashaswimalya Acharjee, Balaraman Ravindran
arXiv:2607. 04978v1 Announce Type: new Abstract: Joint-Embedding Predictive Architectures (JEPAs) underpin a growing family of latent world models for control from raw pixels, but every existing JEPA world model commits at training time to a single inference paradigm: either trajectory optimisation in a learned dynamics model, or direct behaviour cloning.
By Ruslan Rakhimov, George Bredis, Yuriy Maksyuta, Daniil Gavrilov
arXiv:2609.37250v1 Announce Type: cross
Abstract: World-action models (WAMs) couple future visual-state prediction with action generation. By adapting video generators or image-editing models pretrai...
By Yang Zhang, Jiangyuan Zhao, Chenyou Fan, Jiayu Hu, Xiu Yuan, Chenjia Bai, Xiu Li
arXiv:2606. 12808v1 Announce Type: cross Abstract: Adaptive Hamiltonian learning is central to calibrating and characterizing quantum devices.
By Yash Vardhan Tomar, Dheeraj Peddireddy, Vaneet Aggarwal
arXiv:2401. 11512v2 Announce Type: replace-cross Abstract: Identifying the most suitable variables to represent the state is a fundamental challenge in Reinforcement Learning (RL).
By Charles Westphal, Stephen Hailes, Mirco Musolesi
arXiv:2608. 15088v1 Announce Type: cross Abstract: Human-in-the-loop (HIL) online reinforcement learning for real robots must absorb human interventions quickly while continuing to improve beyond the human prior.
By Zihang Wang, Yishan Wang