GRPO-QPS: Target-Preserving Reinforcement Learning for Quantum Posterior Sampling
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
arXiv:2609.14711v1 Announce Type: new Abstract: Reward-based learning can alter the very posterior distribution that scientific inference aims to estimate. GRPO-QM sidesteps this by learning only an...
arXiv:2507. 18606v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) provides a principled framework for decision-making in partially observable environments, which can be modeled as Markov decision processes and compactly represented through dynamic decision Bayesian networks.
Quantum MeanFlow (QMF) is a new quantum generative sampling method that enables single‑step sample generation by learning an average velocity field over a time interval, unlike the multi‑step quantum flow matching (QFM) which requires sequential integration of an ordinary differential equation. Using parameterized quantum circuits, the authors benchmark QMF and QFM on the MNIST dataset, finding that QMF produces lower image quality than multi‑step QFM but outperforms single‑step QFM at every shot count. Both models were executed on IBM quantum computers, and best‑of‑N rejection sampling mitigates device noise without circuit modification, demonstrating QMF’s practicality for efficient single‑step quantum generative sampling.
arXiv:2606. 12808v1 Announce Type: cross Abstract: Adaptive Hamiltonian learning is central to calibrating and characterizing quantum devices.
The paper introduces GenQAS, a tensor network‑guided reinforcement learning framework that uses a learned local transition model to generate synthetic circuit transitions for prioritized generative replay. By mixing these synthetic transitions with real experience during Double Deep Q‑Network updates, GenQAS addresses sample starvation in quantum architecture search. Across benchmarks ranging from 6 to 15 qubits, the method improves success probabilities, identifies compact circuits, and reduces steps to chemical accuracy by up to 92.7%.
Adaptive Hamiltonian learning is central to calibrating and characterizing quantum devices. In an adaptive controller, choosing the next experiment is itself a computation.