arXiv:2507. 22854v3 Announce Type: replace-cross Abstract: We propose novel classical and quantum online algorithms for learning finite- and infinite-horizon Markov Decision Processes (MDPs).
By Andris Ambainis, Joao F. Doriguello, Debbie Lim
arXiv:2507. 18606v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) provides a principled framework for decision-making in partially observable environments, which can be modeled as Markov decision processes and compactly represented through dynamic decision Bayesian networks.
By Gilberto Cunha, Alexandra Ram\^oa, Andr\'e Sequeira, Michael de Oliveira, Lu\'is Barbosa
The paper investigates whether quantum reinforcement learning algorithms can be matched by efficient classical methods. It focuses on a simplified reinforcement learning setting with a uniform generative model, providing finite‑sample guarantees for classical kernelized Fitted Q‑Iteration that uses kernels aligned with parameterized quantum circuits. The authors identify sufficient conditions on data encoding, kernel choice, and problem structure under which this classical approach dequantizes quantum Q‑learning, and suggest using kernelized Fitted Q‑Iteration as a heuristic when those conditions cannot be verified.
By Pablo Rodriguez-Grasa, Sofiene Jerbi, Mikel Sanz, Ryan Sweke
arXiv:2607. 01197v1 Announce Type: new Abstract: Quantum computing has emerged as a promising computational paradigm for machine learning (ML), with the potential to offer computational advantages over classical approaches.
By Chuanming Yu, Jiaming Liu, Zihao Ge, Xiongfei Wu, Lulu Zhu, Pengzhan Zhao, Jianjun Zhao
arXiv:2607. 01080v1 Announce Type: new Abstract: We investigate Gaussian process (GP) bandit optimization with quantum kernels, assuming the mean reward function lies in the reproducing kernel Hilbert space (RKHS) induced by the quantum kernel.
By Yuqi Huang, Vincent Y. F. Tan, Sharu Theresa Jose
Quantum Tiq‑Taq‑Toe is a popular benchmark for quantum computing and machine learning, yet no reinforcement learning (RL) methods have been applied to it. The paper introduces RL techniques for this game, which is simpler than Quantum Chess but still challenging due to partial observability and exponential state complexity. States are represented by a 3×3 measurement matrix and a 9×9 move‑history matrix of entanglement relations, making strategy development difficult because each move can collapse the quantum state.
By Catalin-Viorel Dinu, Thomas Moerland
The paper investigates the use of variational quantum circuits (VQCs) in hierarchical reinforcement learning (HRL). It shows that a hybrid HRL agent incorporating a quantum feature extractor can outperform classical baselines with fewer parameters, but using VQCs for option-value estimation hampers learning. The study also explores how different quantum circuit designs influence performance and proposes design principles for efficient hybrid HRL agents.
By Yu-Ting Lee, Samuel Yen-Chi Chen, Fu-Chieh Chang
arXiv:2607. 21121v1 Announce Type: cross Abstract: In this work, a quantum architecture search framework for approximate quantum state preparation (QSP) is proposed.
By Marco Mordacci, Michele Amoretti
arXiv:2606. 18503v1 Announce Type: new Abstract: Remaining useful life (RUL) estimation is central to predictive maintenance, where an unplanned failure can cost far more than the asset itself.
By Manoranjan Gandhudi, Arunkumar V., G. R. Anil, Gangadharan G. R
The paper introduces BUMEX, a reinforcement learning exploration strategy that leverages a set of prior models containing the true transition kernel and reward function. By optimizing over this model set, the method derives upper and lower bounds on the Q‑function to guide exploration, providing theoretical guarantees of convergence to the optimal policy. When the model set follows a bounded‑parameter MDP structure, the optimization becomes convex, enabling finite‑time convergence under mild assumptions and demonstrating accelerated learning in simulations.
By J. S. van Hulst, W. P. M. H. Heemels, D. J. Antunes
arXiv:2606. 08276v1 Announce Type: cross Abstract: Quantum reinforcement learning (QRL) is a promising approach to learn effective decision strategies across several applications with stochastic environments.
By Alexander DeRieux, Walid Saad
The paper introduces a framework for optimal policy improvement in reinforcement learning, defining it as the best single update under given constraints. It shows that restricting improvement to a subset of states is equivalent to solving an induced Markov Decision Process, linking planning with explicit or implicit models to optimal policy improvement. The authors develop a novel operator for greedification under approximate evaluation, demonstrating empirical gains across several RL algorithms and settings.
By Yaniv Oren, Viliam Vadocz, Wiktor Zabka, Thomas Evers, Jan Robine, Wendelin B\"ohmer, Matthijs T. J. Spaan, Martha White, Hendrik Baier, Fenghui Yu