arXiv:2609.14711v1 Announce Type: new
Abstract: Reward-based learning can alter the very posterior distribution that scientific inference aims to estimate. GRPO-QM sidesteps this by learning only an...
By Yufeng Wang, Parivesh Priye, Lu Wei, Haibin Ling
arXiv:2507. 18606v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) provides a principled framework for decision-making in partially observable environments, which can be modeled as Markov decision processes and compactly represented through dynamic decision Bayesian networks.
By Gilberto Cunha, Alexandra Ram\^oa, Andr\'e Sequeira, Michael de Oliveira, Lu\'is Barbosa
Quantum MeanFlow (QMF) is a new quantum generative sampling method that enables single‑step sample generation by learning an average velocity field over a time interval, unlike the multi‑step quantum flow matching (QFM) which requires sequential integration of an ordinary differential equation. Using parameterized quantum circuits, the authors benchmark QMF and QFM on the MNIST dataset, finding that QMF produces lower image quality than multi‑step QFM but outperforms single‑step QFM at every shot count. Both models were executed on IBM quantum computers, and best‑of‑N rejection sampling mitigates device noise without circuit modification, demonstrating QMF’s practicality for efficient single‑step quantum generative sampling.
By Ashish Joshi, Eshaan Mistry, Takahiko Koyama
arXiv:2606. 12808v1 Announce Type: cross Abstract: Adaptive Hamiltonian learning is central to calibrating and characterizing quantum devices.
By Yash Vardhan Tomar, Dheeraj Peddireddy, Vaneet Aggarwal
The paper introduces GenQAS, a tensor network‑guided reinforcement learning framework that uses a learned local transition model to generate synthetic circuit transitions for prioritized generative replay. By mixing these synthetic transitions with real experience during Double Deep Q‑Network updates, GenQAS addresses sample starvation in quantum architecture search. Across benchmarks ranging from 6 to 15 qubits, the method improves success probabilities, identifies compact circuits, and reduces steps to chemical accuracy by up to 92.7%.
By Akash Kundu, Amit Kumar Jaiswal, Sebastian Feld, Prayag Tiwari
Adaptive Hamiltonian learning is central to calibrating and characterizing quantum devices. In an adaptive controller, choosing the next experiment is itself a computation.
The paper introduces GenQAS, a tensor‑network‑guided reinforcement learning framework that uses a learned local transition model to generate synthetic circuit transitions for prioritized generative replay. By mixing these synthetic transitions with real experience during Double Deep Q‑Network updates, GenQAS addresses sample starvation in quantum architecture search. Experiments on chemical Hamiltonian benchmarks up to 12 qubits and a 15‑qubit Ising model show significant improvements in success probability and circuit compactness, while a noisy 6‑qubit BeH₂ transfer experiment demonstrates a 92.7% reduction in steps to chemical accuracy.
arXiv:2607. 28916v1 Announce Type: cross Abstract: Multistep credit assignment is critical for sample-efficient reinforcement learning, yet managing off-policy bias in Q-learning remains a fundamental challenge.
By Brett Daley
arXiv:2608. 09805v1 Announce Type: cross Abstract: Exploration has been a focus of reinforcement learning research for a long time.
By Vatsal Venkatkrishna, Nico Daheim, Iryna Gurevych
The paper investigates whether quantum reinforcement learning algorithms can be matched by efficient classical methods. It focuses on a simplified reinforcement learning setting with a uniform generative model, providing finite‑sample guarantees for classical kernelized Fitted Q‑Iteration that uses kernels aligned with parameterized quantum circuits. The authors identify sufficient conditions on data encoding, kernel choice, and problem structure under which this classical approach dequantizes quantum Q‑learning, and suggest using kernelized Fitted Q‑Iteration as a heuristic when those conditions cannot be verified.
By Pablo Rodriguez-Grasa, Sofiene Jerbi, Mikel Sanz, Ryan Sweke
arXiv:2608. 19306v1 Announce Type: cross Abstract: Given a set of input states, we consider the task of predicting the expectation value of a Pauli observable at the output of an unknown quantum evolution, using only a limited number of measurements.
By Jonas J\"ager, Yaroslav Khmelnitskiy, Paolo Braccia, Artur Miroszewski, Diego Garc\'ia-Mart\'in, M. Cerezo, Piotr Czarnik
arXiv:2604.21863v2 Announce Type: replace-cross
Abstract: Deep reinforcement learning for quantum circuit optimization faces three bottlenecks: replay buffers that overlook temporal difference (TD) t...
By Akash Kundu, Sebastian Feld