arXiv AI

Aftab: A Comprehensive Benchmark of CNN Encoders and Advanced Value Functions in Parallelized Q-Networks

arXiv:2608. 07335v1 Announce Type: cross Abstract: Recent advancements in deep reinforcement learning have increasingly favored simplified, highly parallelized paradigms.

arXiv AI
Sep 25

Aftab: A Progressive Design Study of Visual Encoders and Value Estimation for Replay-Free Parallelized Q-Learning

The paper presents Aftab, a new architecture for replay‑free parallelized Q‑learning that systematically explores visual encoder designs, multiplicative feature interactions, and value‑estimation strategies. Through a three‑phase study on Atari‑57, the authors compare eight convolutional encoders, integrate Hadamax‑style interactions, and evaluate categorical‑dueling, ensemble‑dueling, and combined configurations, ultimately achieving a higher human‑normalized score than the baseline PQN. Aftab is also evaluated on Procgen Hard, showing improved terminal IQM and learning‑curve area, and the full framework is released as open source.

By Taha Shieenavaz, Shabnam Zareshahraki, Loris Nanni
arXiv AI
6d ago

A Progressive Design Study of Visual Encoders and Value Estimation for Replay-Free Parallelized Q-Learning

The paper presents a three‑phase study of visual encoders and value‑estimation methods for replay‑free parallelized Q‑learning within the Parallelized Q‑Network framework. It compares eight convolutional encoder topologies, adds Hadamax‑style multiplicative interactions and pooling, and evaluates categorical‑dueling, ensemble‑dueling, and categorical ensemble‑dueling value‑estimation configurations. The resulting architecture, Aftab, outperforms a baseline PQN on Atari‑57 and shows improved performance on Procgen Hard, demonstrating that visual topology, multiplicative representation, and downstream value‑estimation design significantly influence replay‑free Q‑learning when considered alongside computational complexity.

By Taha Shieenavaz, Shabnam Zareshahraki, Loris Nanni
arXiv AI
Sep 21

Deep Reinforcement Learning with Buffered Quantile Objectives

The paper introduces Deep-BQRL, a model‑free distributional reinforcement‑learning framework that extends buffered‑quantile learning to neural function approximation. It learns conditional return quantiles from sampled transitions, constructs buffered action scores, and uses ensemble disagreement for exploration, enabling risk‑sensitive decision‑making without explicit return‑law planning. Experiments on asset‑selling and slippery FrozenLake show that Deep‑BQRL achieves smaller mean cumulative point‑quantile policy gaps than PPO and TRPO, while illustrating interpretable risk‑sensitive stopping decisions.

By Mohammad Alipour-vaezi, Sajad Khodadadian
arXiv Machine Learning
Jun 25

Hierarchical Reinforcement Learning for Neural Network Compression (HiReLC): Pruning and Quantization

arXiv:2606. 26002v1 Announce Type: new Abstract: We present HiReLC, a hierarchical ensemble-reinforcement learning framework for automated joint quantization and structured pruning of deep neural networks.

By Kamar Hibatallah Baghdadi, Kawther Guoual Belhamidi, Sara Belhadj, Aissa Boulmerka, Nadir Farhi
arXiv Machine Learning
Jun 10

Discovering Interpretable Multi-Parameter Control Policies for Evolutionary Algorithms Using Deep Reinforcement Learning

arXiv:2606. 10129v1 Announce Type: new Abstract: While deep Reinforcement Learning (deep-RL) has been increasingly applied to parameter control in evolutionary algorithms, rigorous theoretical analysis of parameter control remains largely restricted to single-parameter settings, owing to the difficulty of deriving effective, interpretable multi-parameter policies amenable to formal study.

By Tai Nguyen, Phong Le, Carola Doerr, Nguyen Dang