arXiv Machine Learning

Quantum Hierarchical Reinforcement Learning via Variational Quantum Circuits

The paper investigates the use of variational quantum circuits (VQCs) in hierarchical reinforcement learning (HRL). It shows that a hybrid HRL agent incorporating a quantum feature extractor can outperform classical baselines with fewer parameters, but using VQCs for option-value estimation hampers learning. The study also explores how different quantum circuit designs influence performance and proposes design principles for efficient hybrid HRL agents.

arXiv Machine Learning
Jul 13

Action-Factored Multi-Agent Reinforcement Learning for Scalable Quantum Device Tuning

arXiv:2607. 09422v1 Announce Type: new Abstract: Cooperative multi-agent reinforcement learning is well suited to problems with large parameter spaces and exploitable local structure, such as the tuning of electrostatically-defined quantum-dot arrays.

By Edwin De Nicolo, Rahul Marchand, Cornelius Carlsson, Pranav Vaidhyanathan, Natalia Ares
arXiv AI
Aug 3

DreamQAS: Learning a Decision-Useful World Model for VQE-Efficient Quantum Architecture Search

arXiv:2607. 29491v1 Announce Type: cross Abstract: Reinforcement-learning-based quantum architecture search (RL-QAS) repeatedly optimizes a variational quantum eigensolver (VQE) after extending a circuit, although circuit construction and action legality are deterministic and known.

By Jiayang Niu, Yan Wang, Jie Li, Ke Deng, Azadeh Alavi, Muhammad Usman, Yongli Ren
arXiv AI
Sep 11

Reinforcement learning for Quantum Tiq-Taq-Toe

Quantum Tiq‑Taq‑Toe is a popular benchmark for quantum computing and machine learning, yet no reinforcement learning (RL) methods have been applied to it. The paper introduces RL techniques for this game, which is simpler than Quantum Chess but still challenging due to partial observability and exponential state complexity. States are represented by a 3×3 measurement matrix and a 9×9 move‑history matrix of entanglement relations, making strategy development difficult because each move can collapse the quantum state.

By Catalin-Viorel Dinu, Thomas Moerland
arXiv Machine Learning
Sep 16

Towards Surrogate Based Dequantization of Quantum Reinforcement Learning

The paper investigates whether quantum reinforcement learning algorithms can be matched by efficient classical methods. It focuses on a simplified reinforcement learning setting with a uniform generative model, providing finite‑sample guarantees for classical kernelized Fitted Q‑Iteration that uses kernels aligned with parameterized quantum circuits. The authors identify sufficient conditions on data encoding, kernel choice, and problem structure under which this classical approach dequantizes quantum Q‑learning, and suggest using kernelized Fitted Q‑Iteration as a heuristic when those conditions cannot be verified.

By Pablo Rodriguez-Grasa, Sofiene Jerbi, Mikel Sanz, Ryan Sweke
arXiv AI
Sep 25

Hybrid Variational Quantum-Classical Framework with Adaptive Weighting and Efficiency Assessment

Hybrid Variational Quantum-Classical Framework with Adaptive Weighting and Efficiency Assessment introduces Sim‑HVQC, a hybrid deep quantum neural network that integrates an adaptive, parameter‑free SimAM weighting module with classical feature extraction to retain class‑discriminative information before encoding into a Variational Quantum Circuit. Unlike prior work limited to binary classification, this framework is trained and evaluated on multiple multi‑class datasets such as MNIST, KMNIST, Fashion‑MNIST, and EMNIST. The study highlights reproducibility, parameter efficiency, and interpretability through multi‑seed evaluation, parameter analysis, and latent/quantum feature inspection, with source code publicly available on GitHub.

By Dilli Hang Rai
arXiv AI
Sep 7

A Unified Physics-Aware Quantum Machine Learning Framework across Power GaN HEMTs and Logic Nanowire FETs: Predicting Unseen Process Splits and Held-Out Geometry Combinations with Lower Error and Tighter Split-to-Split Variability

The paper introduces a reinforcement‑learning framework that automatically discovers compact parametrized quantum circuits for modeling power GaN HEMTs and logic nanowire FETs. Using a graph neural network policy trained with proximal policy optimization, the method optimizes circuit architectures based on leave‑one‑group‑out cross‑validation error, achieving the lowest mean absolute error across 11 targets compared to six classical baselines. The results show significant reductions in error and variability for key device metrics (Ioff, VTH, SS) on both HEMT and NWFET datasets, demonstrating the viability of RL‑selected quantum circuits as compact, physically consistent surrogates without explicit physical constraints.

By Rushat Rai, Yun-Yuan Wang, Autsada Kakaen, Pei-Jie Chang, Doan Viet Nguyen, Yuan-Chieh Chiu, Doldet Tantraviwat, Niall Tumilty, Simon See, Wen-Jay Lee, Tai-Yue Li, Nan-Yow Chen, Tian-Li Wu