arXiv Machine Learning

Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems

The paper introduces Adaptive Policy-Guided Error Mitigation (APGEM), a context-aware layer that dynamically selects error mitigation strategies—such as ZNE, PEC, CDR, and REM—during quantum reinforcement learning (QRL) training on NISQ devices. APGEM uses policy-level indicators (quantum-state fidelity, policy entropy, cumulative reward, and approximation ratio) to choose the most suitable mitigation method and integrates it directly into the reinforcement learning loop. Evaluated on the Capacitated Vehicle Routing Problem under various NISQ noise models, APGEM outperforms static mitigation techniques, achieving about 94% of an oracle strategy’s utility, maintaining higher fidelity as noise increases, and producing more stable learning behavior.

arXiv AI
Sep 17

APGEM: Adaptive Policy-Guided Error Mitigation for Quantum Reinforcement Learning on a Real-World CVRP Case Study

The paper introduces APGEM, an adaptive controller that dynamically selects among four error‑mitigation techniques—Zero‑Noise Extrapolation, Probabilistic Error Cancellation, Clifford Data Regression, and Readout Error Mitigation—based on a utility function and Q‑learning scores. Applied to a realistic Delhi‑based Capacitated Vehicle Routing Problem, the adaptive approach improves the quantum reinforcement learning agent’s approximation ratios from 0.84‑0.87 to 0.92‑0.94 under high noise, outperforming constructive heuristics and approaching metaheuristics. The controller’s strategy shifts from a Clifford‑data‑regression‑heavy regime early in training to a balanced use of all techniques as training progresses, demonstrating regime‑dependent selection.

By Shabir Ahmad Sofi, Bisma Majid, Mir Mohammad Yousuf
arXiv AI
Jul 22

Robust Belief-State Policy Learning for Quantum Network Routing Under Decoherence and Time-Varying Conditions

arXiv:2509. 08654v2 Announce Type: replace-cross Abstract: Quantum network routing requires online decisions under probabilistic entanglement generation, finite quantum memories, decoherence, imperfect operations, and classical feedback, while the controller has incomplete knowledge of the physical state.

By Amirhossein Taherpour, Abbas Taherpour, Tamer Khattab, Mazen Hasna
arXiv AI
Aug 3

DreamQAS: Learning a Decision-Useful World Model for VQE-Efficient Quantum Architecture Search

arXiv:2607. 29491v1 Announce Type: cross Abstract: Reinforcement-learning-based quantum architecture search (RL-QAS) repeatedly optimizes a variational quantum eigensolver (VQE) after extending a circuit, although circuit construction and action legality are deterministic and known.

By Jiayang Niu, Yan Wang, Jie Li, Ke Deng, Azadeh Alavi, Muhammad Usman, Yongli Ren
arXiv AI
Sep 24

Quantum Reinforcement Learning for Cost and Delay Tradeoffs in Quantum Cloud Orchestration

Quantum Reinforcement Learning for Cost and Delay Tradeoffs in Quantum Cloud Orchestration proposes QRLQ, a scheduling framework that integrates parameterised quantum circuits with a dueling double deep Q‑network to balance execution cost and delay in quantum‑as‑a‑service environments. Simulation results show QRLQ outperforms heuristic baselines, achieving 5‑11% lower mean cost and up to 82% lower mean delay while maintaining fidelity within 2% of a fidelity‑greedy policy. Compared to a classical deep reinforcement learning baseline, QRLQ delivers comparable performance with 72% fewer trainable parameters.

By An N. H. Phan, Dang Van Huynh, Muhammad Usman, Hoa T. Nguyen
arXiv Machine Learning
Jul 1

Quantum Bayesian Networks Can Speed up Reinforcement Learning in Partially Observable Environments

arXiv:2507. 18606v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) provides a principled framework for decision-making in partially observable environments, which can be modeled as Markov decision processes and compactly represented through dynamic decision Bayesian networks.

By Gilberto Cunha, Alexandra Ram\^oa, Andr\'e Sequeira, Michael de Oliveira, Lu\'is Barbosa
arXiv Machine Learning
Sep 25

Learning and interpreting policies for simultaneous entanglement requests in quantum networks

The paper presents a reinforcement‑learning approach to schedule link‑level entanglement in quantum networks, using a Markov Decision Process and double deep Q‑networks with message‑passing neural networks. The resulting policies achieve 100% success rates even when the link activation probability is reduced by up to 71% compared to baseline heuristics, and maintain at least 80% success when task placements are hardware‑restricted. The authors also develop metrics to interpret the learned policy and employ a large language model to generate a heuristic that matches the DQN performance, suggesting a scalable method for extracting interpretable strategies in large quantum networks.

By Leon Rode, Sumeet Khatri, Supartha Podder