arXiv:2507. 22854v3 Announce Type: replace-cross Abstract: We propose novel classical and quantum online algorithms for learning finite- and infinite-horizon Markov Decision Processes (MDPs).
By Andris Ambainis, Joao F. Doriguello, Debbie Lim
arXiv:2608. 02826v1 Announce Type: cross Abstract: Reinforcement learning is a subfield of machine learning that studies how an agent interacts with an environment in order to extract as large a reward as possible.
By Joao F. Doriguello
arXiv:2607. 29491v1 Announce Type: cross Abstract: Reinforcement-learning-based quantum architecture search (RL-QAS) repeatedly optimizes a variational quantum eigensolver (VQE) after extending a circuit, although circuit construction and action legality are deterministic and known.
By Jiayang Niu, Yan Wang, Jie Li, Ke Deng, Azadeh Alavi, Muhammad Usman, Yongli Ren
arXiv:2606. 08276v1 Announce Type: cross Abstract: Quantum reinforcement learning (QRL) is a promising approach to learn effective decision strategies across several applications with stochastic environments.
By Alexander DeRieux, Walid Saad
arXiv:2607. 01197v1 Announce Type: new Abstract: Quantum computing has emerged as a promising computational paradigm for machine learning (ML), with the potential to offer computational advantages over classical approaches.
By Chuanming Yu, Jiaming Liu, Zihao Ge, Xiongfei Wu, Lulu Zhu, Pengzhan Zhao, Jianjun Zhao
Quantum Tiq‑Taq‑Toe is a popular benchmark for quantum computing and machine learning, yet no reinforcement learning (RL) methods have been applied to it. The paper introduces RL techniques for this game, which is simpler than Quantum Chess but still challenging due to partial observability and exponential state complexity. States are represented by a 3×3 measurement matrix and a 9×9 move‑history matrix of entanglement relations, making strategy development difficult because each move can collapse the quantum state.
By Catalin-Viorel Dinu, Thomas Moerland
arXiv:2609.00455v1 Announce Type: new
Abstract: Large language models (LLMs) are being used as policies for autonomous decision-making and planning in many domains. Despite their strong reasoning cap...
By Shubham Kumar, Harshit Kumar, Narendra Ahuja, Saurabh Jha
arXiv:2507. 18606v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) provides a principled framework for decision-making in partially observable environments, which can be modeled as Markov decision processes and compactly represented through dynamic decision Bayesian networks.
By Gilberto Cunha, Alexandra Ram\^oa, Andr\'e Sequeira, Michael de Oliveira, Lu\'is Barbosa
QART is a quantum‑classical hybrid architecture that augments a language model with quantum encoding, CIM‑based QUBO optimization, and quantum decoding to improve long‑horizon reasoning. The authors claim that, under certain assumptions, QART can maintain a non‑zero probability of recovering an optimal reasoning path while traditional autoregressive LLMs see their acceptance probability drop to zero as cumulative risk grows. Experiments on six benchmarks with three backbone models show that QART outperforms the baselines in 14 of 15 pairings, with relative gains up to 84.0% on SciCode.
By Lehao Lin, Yuheng Cheng, Guolong Liu, Yao Li, Xuning Tan, Xiyuan Zhou, Ruixi Zou, Shi Wang, Huan Zhao, Wenxuan Liu, Haifeng Wu, Junhua Zhao
The paper investigates whether quantum reinforcement learning algorithms can be matched by efficient classical methods. It focuses on a simplified reinforcement learning setting with a uniform generative model, providing finite‑sample guarantees for classical kernelized Fitted Q‑Iteration that uses kernels aligned with parameterized quantum circuits. The authors identify sufficient conditions on data encoding, kernel choice, and problem structure under which this classical approach dequantizes quantum Q‑learning, and suggest using kernelized Fitted Q‑Iteration as a heuristic when those conditions cannot be verified.
By Pablo Rodriguez-Grasa, Sofiene Jerbi, Mikel Sanz, Ryan Sweke
arXiv:2603. 10289v2 Announce Type: replace-cross Abstract: Whether uniquely quantum resources confer advantages in fully classical, competitive environments remains an open question.
By Peiyong Wang, Kieran Hymas, James Quach
arXiv:2608. 05371v1 Announce Type: new Abstract: World models learn latent states that summarize interaction histories, evolve over time, and support prediction, simulation, or planning.
By Hailong Jiang, Emran Hossain, Feng Yu, Jianfeng Zhu, Guilin Zhang, Wulan Guo