Average-and Last-Iterate Lower Bounds for Optimistic Matrix Mirror-Prox in Quantum Zero-Sum Games
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
arXiv:2606. 02655v1 Announce Type: cross Abstract: External regret certifies stability only against replacing one's behavior by a fixed alternative.
arXiv:2610. 01181v1 Announce Type: new Abstract: We consider stochastic games with independent controlled chains and unknown transition kernels, where players observe only their local states and realized payoffs.
arXiv:2608. 04149v1 Announce Type: cross Abstract: Swap regret governs the rate at which uncoupled learning dynamics converge to correlated equilibria in multiplayer general-sum games.
arXiv:2609.38375v1 Announce Type: new Abstract: Can a constant number of linear minimizations per round improve on the $T^{3/4}$ regret rate of online Frank-Wolfe on general convex sets? Weibel et al...
arXiv:2609.22839v1 Announce Type: cross Abstract: Can simple learning rules keep their regret bounded in self-play? Recent work achieves constant regret bounds through modified regularization and hig...
arXiv:2608. 14319v1 Announce Type: new Abstract: We study quantum multi-armed bandits (QMAB) and quantum linear bandits (QLB) in the model of Wan et al.