arXiv Machine Learning

Variance-Reduced Q-Learning over Static and Time-Varying Networks

arXiv:2607. 21876v1 Announce Type: new Abstract: We investigate a decentralized reinforcement learning problem involving multiple agents that interact with the same Markov Decision Process (MDP).

arXiv Machine Learning
Sep 18

Robust Federated Q-Learning with Almost No Communication

The paper introduces Robust Fed-Q, a federated Q‑learning algorithm designed for settings where multiple agents interact with a shared Markov Decision Process and communicate through a central server. It combines model‑based and model‑free reinforcement learning techniques with a median‑of‑means strategy from robust statistics to handle a small fraction of adversarial agents. The authors prove that Robust Fed-Q achieves exact convergence to the optimal value function with high probability, attains near‑optimal finite‑time rates that benefit from collaboration, and requires only “~O(1)” communication rounds per guarantee.

By Sreejeet Maity, Aritra Mitra
arXiv AI
Sep 17

Decentralized Optimal Equilibrium Learning Over Dynamic Networks

The paper introduces a decentralized learning framework for finding socially optimal equilibria in finite normal-form games played over dynamic communication networks. Agents only observe their own payoffs, lack prior knowledge of the game, and communicate with time-varying neighbors using low-bandwidth, time-stamped tables instead of raw actions or payoff data. The proposed dynamics combine randomized semantic signals, table fusion, and temporal majority reconstruction to achieve finite-time logarithmic regret guarantees for optimal equilibrium selection under utilitarian and proportional-fair social welfare objectives, as demonstrated by simulations.

By Seref Taha Kiremitci, Muhammed O. Sayin
arXiv AI
Jul 17

Reinforcement Learning in Switching Non-Stationary Markov Decision Processes: Algorithms and Convergence Analysis

arXiv:2503. 18607v2 Announce Type: replace-cross Abstract: We introduce the Switching Non-Stationary Markov Decision Process (SNS-MDP) framework, in which the environment transitions among a finite set of MDPs governed by a latent Markov chain while the agent observes only the external state.

By Mohsen Amiri, Sindri Magn\'usson
arXiv AI
Sep 2

Provably Efficient Federated Reinforcement Learning with Linear Function Approximation and Logarithmic Communication Cost

The paper introduces Fed‑LSVI, a federated online reinforcement learning algorithm that uses linear function approximation in episodic Markov decision processes. It achieves a regret bound of ≥O(√{Md^3H^4T}) while only exchanging compressed sufficient statistics, thereby meeting privacy constraints. The method reduces communication cost to logarithmic in the number of episodes, a marked improvement over previous approaches that required linear communication.

By Zihang Liang, Haochen Zhang, Lingzhou Xue
arXiv Machine Learning
Sep 23

A Decentralized Partially Observable Team Decision Methodology with Delayed Information Sharing

The paper introduces a decentralized decision-making framework for teams operating under partial observability and unknown system dynamics. By leveraging low-rank latent dynamics and delayed shared information, each team member learns an approximate Markov decision process using only local private data and delayed common updates. The resulting algorithm achieves near‑optimal team performance without requiring a centralized coordinator or training, and the authors provide finite‑sample guarantees and a sample‑complexity bound.

By Xiaoxing Ren, Thomas Parisini, Andreas A. Malikopoulos