The paper introduces Robust Fed-Q, a federated Q‑learning algorithm designed for settings where multiple agents interact with a shared Markov Decision Process and communicate through a central server. It combines model‑based and model‑free reinforcement learning techniques with a median‑of‑means strategy from robust statistics to handle a small fraction of adversarial agents. The authors prove that Robust Fed-Q achieves exact convergence to the optimal value function with high probability, attains near‑optimal finite‑time rates that benefit from collaboration, and requires only “~O(1)” communication rounds per guarantee.
By Sreejeet Maity, Aritra Mitra
The paper introduces a decentralized learning framework for finding socially optimal equilibria in finite normal-form games played over dynamic communication networks. Agents only observe their own payoffs, lack prior knowledge of the game, and communicate with time-varying neighbors using low-bandwidth, time-stamped tables instead of raw actions or payoff data. The proposed dynamics combine randomized semantic signals, table fusion, and temporal majority reconstruction to achieve finite-time logarithmic regret guarantees for optimal equilibrium selection under utilitarian and proportional-fair social welfare objectives, as demonstrated by simulations.
By Seref Taha Kiremitci, Muhammed O. Sayin
arXiv:2606. 00266v1 Announce Type: cross Abstract: A long-standing challenge in distributed wireless systems is ensuring efficient and fair random channel access.
By Kamil Szczech, Maksymilian Wojnar, Krzysztof Rusek, Katarzyna Kosek-Szott, Szymon Szott
arXiv:2503. 18607v2 Announce Type: replace-cross Abstract: We introduce the Switching Non-Stationary Markov Decision Process (SNS-MDP) framework, in which the environment transitions among a finite set of MDPs governed by a latent Markov chain while the agent observes only the external state.
By Mohsen Amiri, Sindri Magn\'usson
The paper introduces Fed‑LSVI, a federated online reinforcement learning algorithm that uses linear function approximation in episodic Markov decision processes. It achieves a regret bound of ≥O(√{Md^3H^4T}) while only exchanging compressed sufficient statistics, thereby meeting privacy constraints. The method reduces communication cost to logarithmic in the number of episodes, a marked improvement over previous approaches that required linear communication.
By Zihang Liang, Haochen Zhang, Lingzhou Xue
The paper introduces a decentralized decision-making framework for teams operating under partial observability and unknown system dynamics. By leveraging low-rank latent dynamics and delayed shared information, each team member learns an approximate Markov decision process using only local private data and delayed common updates. The resulting algorithm achieves near‑optimal team performance without requiring a centralized coordinator or training, and the authors provide finite‑sample guarantees and a sample‑complexity bound.
By Xiaoxing Ren, Thomas Parisini, Andreas A. Malikopoulos