arXiv:2608. 07335v1 Announce Type: cross Abstract: Recent advancements in deep reinforcement learning have increasingly favored simplified, highly parallelized paradigms.
By Taha Shieenavaz, Shabnam Zareshahraki, Loris Nanni
arXiv:2608. 03069v1 Announce Type: new Abstract: Deep Q-Networks (DQNs) learn value functions through bootstrapped temporal-difference updates, where future returns are approximated using a greedy maximization over next-state action values.
By Lipeng Zu, Xiaonan Zhang
The paper presents Aftab, a new architecture for replay‑free parallelized Q‑learning that systematically explores visual encoder designs, multiplicative feature interactions, and value‑estimation strategies. Through a three‑phase study on Atari‑57, the authors compare eight convolutional encoders, integrate Hadamax‑style interactions, and evaluate categorical‑dueling, ensemble‑dueling, and combined configurations, ultimately achieving a higher human‑normalized score than the baseline PQN. Aftab is also evaluated on Procgen Hard, showing improved terminal IQM and learning‑curve area, and the full framework is released as open source.
By Taha Shieenavaz, Shabnam Zareshahraki, Loris Nanni
The paper presents a three‑phase study of visual encoders and value‑estimation methods for replay‑free parallelized Q‑learning within the Parallelized Q‑Network framework. It compares eight convolutional encoder topologies, adds Hadamax‑style multiplicative interactions and pooling, and evaluates categorical‑dueling, ensemble‑dueling, and categorical ensemble‑dueling value‑estimation configurations. The resulting architecture, Aftab, outperforms a baseline PQN on Atari‑57 and shows improved performance on Procgen Hard, demonstrating that visual topology, multiplicative representation, and downstream value‑estimation design significantly influence replay‑free Q‑learning when considered alongside computational complexity.
By Taha Shieenavaz, Shabnam Zareshahraki, Loris Nanni
arXiv:2609.06421v1 Announce Type: cross
Abstract: Batch normalization (BN) substantially improves sample efficiency in continuous-control actor-critic methods such as CrossQ, yet recent studies repor...
By Daniel Palenicek, Mikael Henaff, Scott Fujimoto, Koustuv Sinha
arXiv:2606. 29806v1 Announce Type: cross Abstract: Action-values are foundational to many control algorithms such as Q-learning.
By Prabhat Nagarajan, Brett Daley, Martha White, Marlos C. Machado
arXiv:2607. 27203v1 Announce Type: new Abstract: Pre-training followed by fine-tuning has become the dominant recipe for learning performant policies, and in value-based reinforcement learning (RL) this raises a natural question: given a pretrained policy, should the Q-function be pretrained on offline data too?
By Perry Dong, Ron Polonsky, Dorsa Sadigh, Chelsea Fin
arXiv:2511. 03836v2 Announce Type: replace Abstract: Deep Q-Networks (DQNs) estimate future returns by learning from transitions sampled from a replay buffer.
By Lipeng Zu, Hansong Zhou, Xiaonan Zhang
Pre-training followed by fine-tuning has become the dominant recipe for learning performant policies, and in value-based reinforcement learning (RL) this raises a natural question: given a pretrained policy, should the Q-function be pretrained on offline data too? Conventional wisdom suggests it should, but recent results show that online RL with a randomly-initialized Q-function can result in highly performant and reliable policies without needing to pretrain the Q-function.
arXiv:2607. 18722v1 Announce Type: new Abstract: Asynchronous reinforcement learning improves throughput by decoupling rollout generation from optimization, but staleness is an inevitable byproduct compounded by policy lag, engine delays, and mixture-of-experts routing.
By Junyao Yang, Yucheng Shi, Zongxia Li, Zhongzhi Li, Ruhan Wang, Xiangxin Zhou, Kishan Panaganti, Haitao Mi, Leowei Liang
arXiv:2602. 03120v2 Announce Type: replace-cross Abstract: Post-Training Quantization (PTQ) is essential for deploying Large Language Models (LLMs) on memory-constrained devices, yet it renders models static and difficult to fine-tune.
By Yinggan Xu, Kajetan Schweighofer, Risto Miikkulainen, Xin Qiu
arXiv:2506. 05716v2 Announce Type: replace-cross Abstract: Deep Q-Networks (DQN) can suffer from overestimation bias because bootstrapped targets use a maximisation operation over noisy value estimates.
By Adrian Ly, Richard Dazeley, Peter Vamplew, Francisco Cruz, Sunil Aryal