arXiv:2607. 15440v1 Announce Type: new Abstract: We introduce Stochastic Reset Pathfinding (SRP), an episodic learning problem on a known directed graph with unknown stationary edge success probabilities.
By Guni Sharon, Wei Zhang
The paper presents a reinforcement‑learning approach to schedule link‑level entanglement in quantum networks, using a Markov Decision Process and double deep Q‑networks with message‑passing neural networks. The resulting policies achieve 100% success rates even when the link activation probability is reduced by up to 71% compared to baseline heuristics, and maintain at least 80% success when task placements are hardware‑restricted. The authors also develop metrics to interpret the learned policy and employ a large language model to generate a heuristic that matches the DQN performance, suggesting a scalable method for extracting interpretable strategies in large quantum networks.
By Leon Rode, Sumeet Khatri, Supartha Podder
The paper introduces GenQAS, a tensor network‑guided reinforcement learning framework that uses a learned local transition model to generate synthetic circuit transitions for prioritized generative replay. By mixing these synthetic transitions with real experience during Double Deep Q‑Network updates, GenQAS addresses sample starvation in quantum architecture search. Across benchmarks ranging from 6 to 15 qubits, the method improves success probabilities, identifies compact circuits, and reduces steps to chemical accuracy by up to 92.7%.
By Akash Kundu, Amit Kumar Jaiswal, Sebastian Feld, Prayag Tiwari
arXiv:2607. 06230v1 Announce Type: cross Abstract: Parameterized quantum circuits (PQCs) are increasingly used as policies and value functions in quantum reinforcement learning, yet it remains unclear when and why quantum policies generalize.
By Jian Xu, Delu Zeng, John Paisley, Qibin Zhao
The paper introduces GenQAS, a tensor‑network‑guided reinforcement learning framework that uses a learned local transition model to generate synthetic circuit transitions for prioritized generative replay. By mixing these synthetic transitions with real experience during Double Deep Q‑Network updates, GenQAS addresses sample starvation in quantum architecture search. Experiments on chemical Hamiltonian benchmarks up to 12 qubits and a 15‑qubit Ising model show significant improvements in success probability and circuit compactness, while a noisy 6‑qubit BeH₂ transfer experiment demonstrates a 92.7% reduction in steps to chemical accuracy.
arXiv:2609.05842v1 Announce Type: cross
Abstract: Reinforcement learning with verifiable rewards enables large language models to think slowly, but the same training can induce policy collapse: proba...
By Xiansheng Cai, Xiu-Hao Deng, Kun Chen