arXiv:2607. 15440v1 Announce Type: new Abstract: We introduce Stochastic Reset Pathfinding (SRP), an episodic learning problem on a known directed graph with unknown stationary edge success probabilities.
By Guni Sharon, Wei Zhang
The paper presents a reinforcement‑learning approach to schedule link‑level entanglement in quantum networks, using a Markov Decision Process and double deep Q‑networks with message‑passing neural networks. The resulting policies achieve 100% success rates even when the link activation probability is reduced by up to 71% compared to baseline heuristics, and maintain at least 80% success when task placements are hardware‑restricted. The authors also develop metrics to interpret the learned policy and employ a large language model to generate a heuristic that matches the DQN performance, suggesting a scalable method for extracting interpretable strategies in large quantum networks.
By Leon Rode, Sumeet Khatri, Supartha Podder
The paper introduces GenQAS, a tensor network‑guided reinforcement learning framework that uses a learned local transition model to generate synthetic circuit transitions for prioritized generative replay. By mixing these synthetic transitions with real experience during Double Deep Q‑Network updates, GenQAS addresses sample starvation in quantum architecture search. Across benchmarks ranging from 6 to 15 qubits, the method improves success probabilities, identifies compact circuits, and reduces steps to chemical accuracy by up to 92.7%.
By Akash Kundu, Amit Kumar Jaiswal, Sebastian Feld, Prayag Tiwari
arXiv:2607. 06230v1 Announce Type: cross Abstract: Parameterized quantum circuits (PQCs) are increasingly used as policies and value functions in quantum reinforcement learning, yet it remains unclear when and why quantum policies generalize.
By Jian Xu, Delu Zeng, John Paisley, Qibin Zhao
The paper introduces GenQAS, a tensor‑network‑guided reinforcement learning framework that uses a learned local transition model to generate synthetic circuit transitions for prioritized generative replay. By mixing these synthetic transitions with real experience during Double Deep Q‑Network updates, GenQAS addresses sample starvation in quantum architecture search. Experiments on chemical Hamiltonian benchmarks up to 12 qubits and a 15‑qubit Ising model show significant improvements in success probability and circuit compactness, while a noisy 6‑qubit BeH₂ transfer experiment demonstrates a 92.7% reduction in steps to chemical accuracy.
arXiv:2609.05842v1 Announce Type: cross
Abstract: Reinforcement learning with verifiable rewards enables large language models to think slowly, but the same training can induce policy collapse: proba...
By Xiansheng Cai, Xiu-Hao Deng, Kun Chen
arXiv:2603. 10289v2 Announce Type: replace-cross Abstract: Whether uniquely quantum resources confer advantages in fully classical, competitive environments remains an open question.
By Peiyong Wang, Kieran Hymas, James Quach
arXiv:2606. 31536v1 Announce Type: new Abstract: As Quantum Machine Learning (QML) transitions toward practical implementation, the field faces a critical architectural bottleneck that challenges the fundamental assumptions of classical statistical learning theory.
By Kung-Ming Lan
arXiv:2608. 14319v1 Announce Type: new Abstract: We study quantum multi-armed bandits (QMAB) and quantum linear bandits (QLB) in the model of Wan et al.
By Maoli Liu, Zhuohua Li, John C. S. Lui
QART is a quantum‑classical hybrid architecture that augments a language model with quantum encoding, CIM‑based QUBO optimization, and quantum decoding to improve long‑horizon reasoning. The authors claim that, under certain assumptions, QART can maintain a non‑zero probability of recovering an optimal reasoning path while traditional autoregressive LLMs see their acceptance probability drop to zero as cumulative risk grows. Experiments on six benchmarks with three backbone models show that QART outperforms the baselines in 14 of 15 pairings, with relative gains up to 84.0% on SciCode.
By Lehao Lin, Yuheng Cheng, Guolong Liu, Yao Li, Xuning Tan, Xiyuan Zhou, Ruixi Zou, Shi Wang, Huan Zhao, Wenxuan Liu, Haifeng Wu, Junhua Zhao
arXiv:2507. 18606v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) provides a principled framework for decision-making in partially observable environments, which can be modeled as Markov decision processes and compactly represented through dynamic decision Bayesian networks.
By Gilberto Cunha, Alexandra Ram\^oa, Andr\'e Sequeira, Michael de Oliveira, Lu\'is Barbosa
arXiv:2507. 22854v3 Announce Type: replace-cross Abstract: We propose novel classical and quantum online algorithms for learning finite- and infinite-horizon Markov Decision Processes (MDPs).
By Andris Ambainis, Joao F. Doriguello, Debbie Lim