arXiv:2607. 15440v1 Announce Type: new Abstract: We introduce Stochastic Reset Pathfinding (SRP), an episodic learning problem on a known directed graph with unknown stationary edge success probabilities.
By Guni Sharon, Wei Zhang
arXiv:2607. 06230v1 Announce Type: cross Abstract: Parameterized quantum circuits (PQCs) are increasingly used as policies and value functions in quantum reinforcement learning, yet it remains unclear when and why quantum policies generalize.
By Jian Xu, Delu Zeng, John Paisley, Qibin Zhao
arXiv:2603. 10289v2 Announce Type: replace-cross Abstract: Whether uniquely quantum resources confer advantages in fully classical, competitive environments remains an open question.
By Peiyong Wang, Kieran Hymas, James Quach
arXiv:2606. 31536v1 Announce Type: new Abstract: As Quantum Machine Learning (QML) transitions toward practical implementation, the field faces a critical architectural bottleneck that challenges the fundamental assumptions of classical statistical learning theory.
By Kung-Ming Lan
arXiv:2608. 14319v1 Announce Type: new Abstract: We study quantum multi-armed bandits (QMAB) and quantum linear bandits (QLB) in the model of Wan et al.
By Maoli Liu, Zhuohua Li, John C. S. Lui
arXiv:2507. 18606v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) provides a principled framework for decision-making in partially observable environments, which can be modeled as Markov decision processes and compactly represented through dynamic decision Bayesian networks.
By Gilberto Cunha, Alexandra Ram\^oa, Andr\'e Sequeira, Michael de Oliveira, Lu\'is Barbosa
arXiv:2507. 22854v3 Announce Type: replace-cross Abstract: We propose novel classical and quantum online algorithms for learning finite- and infinite-horizon Markov Decision Processes (MDPs).
By Andris Ambainis, Joao F. Doriguello, Debbie Lim
arXiv:2607. 13988v1 Announce Type: new Abstract: Multi-turn agents solve complex tasks through extended sequences of tool interactions before producing a final answer, making credit assignment a fundamental challenge during post-training.
By Leitian Tao, Baolin Peng, Wenlin Yao, Tao Ge, Hao Cheng, Mike Hang Wang, Jianfeng Gao, Sharon Li
arXiv:2607. 01080v1 Announce Type: new Abstract: We investigate Gaussian process (GP) bandit optimization with quantum kernels, assuming the mean reward function lies in the reproducing kernel Hilbert space (RKHS) induced by the quantum kernel.
By Yuqi Huang, Vincent Y. F. Tan, Sharu Theresa Jose
arXiv:2606. 14929v1 Announce Type: cross Abstract: Modern recommendation systems increasingly rely on dynamically routing diverse queries to multiple embedding models.
By Yan Dai, Negin Golrezaei, Patrick Jaillet
arXiv:2606. 10448v1 Announce Type: cross Abstract: The financial market is a typical low signal-to-noise ratio (SNR) setting, which often destabilizes off-policy maximum-entropy methods like Soft Actor-Critic (SAC).
By Zeyu Liu, Xuanzhi Feng, Sing Kwong Lai, Yuanchen Gao, Xiaoyi Pang, Hualei Zhang, Jingcai Guo, Jie Zhang, Song Guo
arXiv:2606. 12816v2 Announce Type: replace-cross Abstract: Quantum circuit routing is a key step in compiling programs for noisy intermediate-scale quantum processors.
By Yash Vardhan Tomar, Dheeraj Peddireddy