arXiv AI

Riemannian Optimization for Multi-Player Quantum Games on Product Unitary Manifolds

The paper extends the Eisert-Wilkens-Lewenstein quantum game to multiplayer settings with mixed strategies, where each player selects unitary operators and mixes them classically. It introduces the Unitary Strategy Matrix Exponential Algorithm (USMEA), a geometry-aware sequential method that jointly learns local unitary actions and mixing probabilities for each player. The authors analyze USMEA’s convergence under standard conditions and confirm the theory with numerical experiments, demonstrating how classical optimization can be integrated into engineered quantum strategic interactions.

arXiv AI
Sep 11

Reinforcement learning for Quantum Tiq-Taq-Toe

Quantum Tiq‑Taq‑Toe is a popular benchmark for quantum computing and machine learning, yet no reinforcement learning (RL) methods have been applied to it. The paper introduces RL techniques for this game, which is simpler than Quantum Chess but still challenging due to partial observability and exponential state complexity. States are represented by a 3×3 measurement matrix and a 9×9 move‑history matrix of entanglement relations, making strategy development difficult because each move can collapse the quantum state.

By Catalin-Viorel Dinu, Thomas Moerland
arXiv Machine Learning
Jul 13

Action-Factored Multi-Agent Reinforcement Learning for Scalable Quantum Device Tuning

arXiv:2607. 09422v1 Announce Type: new Abstract: Cooperative multi-agent reinforcement learning is well suited to problems with large parameter spaces and exploitable local structure, such as the tuning of electrostatically-defined quantum-dot arrays.

By Edwin De Nicolo, Rahul Marchand, Cornelius Carlsson, Pranav Vaidhyanathan, Natalia Ares
arXiv Machine Learning
Sep 16

Towards Surrogate Based Dequantization of Quantum Reinforcement Learning

The paper investigates whether quantum reinforcement learning algorithms can be matched by efficient classical methods. It focuses on a simplified reinforcement learning setting with a uniform generative model, providing finite‑sample guarantees for classical kernelized Fitted Q‑Iteration that uses kernels aligned with parameterized quantum circuits. The authors identify sufficient conditions on data encoding, kernel choice, and problem structure under which this classical approach dequantizes quantum Q‑learning, and suggest using kernelized Fitted Q‑Iteration as a heuristic when those conditions cannot be verified.

By Pablo Rodriguez-Grasa, Sofiene Jerbi, Mikel Sanz, Ryan Sweke
arXiv AI
Jul 22

Robust Belief-State Policy Learning for Quantum Network Routing Under Decoherence and Time-Varying Conditions

arXiv:2509. 08654v2 Announce Type: replace-cross Abstract: Quantum network routing requires online decisions under probabilistic entanglement generation, finite quantum memories, decoherence, imperfect operations, and classical feedback, while the controller has incomplete knowledge of the physical state.

By Amirhossein Taherpour, Abbas Taherpour, Tamer Khattab, Mazen Hasna
arXiv Machine Learning
Jun 5

DNQ: Deep Nash Q-Network for Partially Observable n-Player Games

arXiv:2606. 06480v1 Announce Type: cross Abstract: Many real-world competitive systems require multiple decision-makers to act simultaneously under shared constraints, limited information, and repeated interaction, as in auctions, resource allocation, and security competition.

By Qintong Xie, Edward Koh, Xavier Cadet, Peter Chin
arXiv Machine Learning
Sep 25

Learning and interpreting policies for simultaneous entanglement requests in quantum networks

The paper presents a reinforcement‑learning approach to schedule link‑level entanglement in quantum networks, using a Markov Decision Process and double deep Q‑networks with message‑passing neural networks. The resulting policies achieve 100% success rates even when the link activation probability is reduced by up to 71% compared to baseline heuristics, and maintain at least 80% success when task placements are hardware‑restricted. The authors also develop metrics to interpret the learned policy and employ a large language model to generate a heuristic that matches the DQN performance, suggesting a scalable method for extracting interpretable strategies in large quantum networks.

By Leon Rode, Sumeet Khatri, Supartha Podder