The paper investigates the use of variational quantum circuits (VQCs) in hierarchical reinforcement learning (HRL). It shows that a hybrid HRL agent incorporating a quantum feature extractor can outperform classical baselines with fewer parameters, but using VQCs for option-value estimation hampers learning. The study also explores how different quantum circuit designs influence performance and proposes design principles for efficient hybrid HRL agents.
By Yu-Ting Lee, Samuel Yen-Chi Chen, Fu-Chieh Chang
arXiv:2607. 21121v1 Announce Type: cross Abstract: In this work, a quantum architecture search framework for approximate quantum state preparation (QSP) is proposed.
By Marco Mordacci, Michele Amoretti
Quantum Tiq‑Taq‑Toe is a popular benchmark for quantum computing and machine learning, yet no reinforcement learning (RL) methods have been applied to it. The paper introduces RL techniques for this game, which is simpler than Quantum Chess but still challenging due to partial observability and exponential state complexity. States are represented by a 3×3 measurement matrix and a 9×9 move‑history matrix of entanglement relations, making strategy development difficult because each move can collapse the quantum state.
By Catalin-Viorel Dinu, Thomas Moerland
The paper investigates whether quantum reinforcement learning algorithms can be matched by efficient classical methods. It focuses on a simplified reinforcement learning setting with a uniform generative model, providing finite‑sample guarantees for classical kernelized Fitted Q‑Iteration that uses kernels aligned with parameterized quantum circuits. The authors identify sufficient conditions on data encoding, kernel choice, and problem structure under which this classical approach dequantizes quantum Q‑learning, and suggest using kernelized Fitted Q‑Iteration as a heuristic when those conditions cannot be verified.
By Pablo Rodriguez-Grasa, Sofiene Jerbi, Mikel Sanz, Ryan Sweke
arXiv:2607. 00365v1 Announce Type: cross Abstract: Artificial intelligence (AI) and quantum information (QI) are rapidly co-evolving.
By Min Chen, Yu Gan, Xin Jin, Yuqing Li, Junqi Wang, Zeguan Wu, Yunfei Wang, Bingzhi Zhang, Priyam Srivastava, Tianlong Chen, Ankit Kulshrestha, Yuan Liu, Juan Jos\'e Mendoza-Arenas, Kaushik P. Seshadreesan, Sarvagya Upadhyay, Xueyue Zhang, Quntao Zhuang, Junyu Liu
arXiv:2603. 10289v2 Announce Type: replace-cross Abstract: Whether uniquely quantum resources confer advantages in fully classical, competitive environments remains an open question.
By Peiyong Wang, Kieran Hymas, James Quach
arXiv:2608. 02826v1 Announce Type: cross Abstract: Reinforcement learning is a subfield of machine learning that studies how an agent interacts with an environment in order to extract as large a reward as possible.
By Joao F. Doriguello
arXiv:2406. 07884v3 Announce Type: replace-cross Abstract: Using partial knowledge of a quantum state to control multiqubit entanglement is a largely unexplored paradigm in the emerging field of quantum interactive dynamics with the potential to address outstanding challenges in quantum state preparation and compression, quantum control, and quantum complexity.
By Pavel Tashev, Stefan Petrov, Matthew T. Diaz, Friederike Metz, Alaina M. Green, Norbert M. Linke, Marin Bukov
The paper presents a reinforcement‑learning approach to schedule link‑level entanglement in quantum networks, using a Markov Decision Process and double deep Q‑networks with message‑passing neural networks. The resulting policies achieve 100% success rates even when the link activation probability is reduced by up to 71% compared to baseline heuristics, and maintain at least 80% success when task placements are hardware‑restricted. The authors also develop metrics to interpret the learned policy and employ a large language model to generate a heuristic that matches the DQN performance, suggesting a scalable method for extracting interpretable strategies in large quantum networks.
By Leon Rode, Sumeet Khatri, Supartha Podder
Neural quantum states (NQS) provide a flexible and scalable framework for approximating quantum many-body wavefunctions. Among NQS parameterizations, autoregressive models are especially attractive because they enable exact, independent sampling from the Born distribution, avoiding the autocorrelation and mixing issues of Markov chain methods.
The paper introduces QFWP-ANO, a quantum neural network architecture that uses a classical hypernetwork to program variational quantum circuit parameters and non‑local observables conditioned on each input. Unlike existing adaptive non‑local observable (ANO) methods that learn a single static observable, QFWP-ANO dynamically adapts to each input. Experiments on multivariate time‑series forecasting and reinforcement learning tasks show that QFWP-ANO outperforms traditional ANO‑based VQCs and other strong baselines, achieving the lowest mean‑squared error in most settings.
By Yu-Ting Lee, Samuel Yen-Chi Chen, Huan-Hsin Tseng
arXiv:2511. 12482v2 Announce Type: replace-cross Abstract: Quantum error correction is essential for fault-tolerant quantum computing.
By Yue Yin, Tailong Xiao, Xiaoyang Deng, Ming He, Jianping Fan, Guihua Zeng