arXiv Machine Learning
Jul 1

Quantum Bayesian Networks Can Speed up Reinforcement Learning in Partially Observable Environments

arXiv:2507. 18606v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) provides a principled framework for decision-making in partially observable environments, which can be modeled as Markov decision processes and compactly represented through dynamic decision Bayesian networks.

By Gilberto Cunha, Alexandra Ram\^oa, Andr\'e Sequeira, Michael de Oliveira, Lu\'is Barbosa
arXiv Machine Learning
Sep 3

Quantum MeanFlow: single-shot generative sampling on NISQ hardware

Quantum MeanFlow (QMF) is a new quantum generative sampling method that enables single‑step sample generation by learning an average velocity field over a time interval, unlike the multi‑step quantum flow matching (QFM) which requires sequential integration of an ordinary differential equation. Using parameterized quantum circuits, the authors benchmark QMF and QFM on the MNIST dataset, finding that QMF produces lower image quality than multi‑step QFM but outperforms single‑step QFM at every shot count. Both models were executed on IBM quantum computers, and best‑of‑N rejection sampling mitigates device noise without circuit modification, demonstrating QMF’s practicality for efficient single‑step quantum generative sampling.

By Ashish Joshi, Eshaan Mistry, Takahiko Koyama
arXiv Machine Learning
Sep 11

Generative Replay Mitigates Sample Starvation in Quantum Architecture Search

The paper introduces GenQAS, a tensor network‑guided reinforcement learning framework that uses a learned local transition model to generate synthetic circuit transitions for prioritized generative replay. By mixing these synthetic transitions with real experience during Double Deep Q‑Network updates, GenQAS addresses sample starvation in quantum architecture search. Across benchmarks ranging from 6 to 15 qubits, the method improves success probabilities, identifies compact circuits, and reduces steps to chemical accuracy by up to 92.7%.

By Akash Kundu, Amit Kumar Jaiswal, Sebastian Feld, Prayag Tiwari