arXiv:2607. 11075v1 Announce Type: new Abstract: The choice of Modulation and Coding (MCS) type for a particular channel condition is made through link adaptation (LA) algorithms that operate at the MAC layer.
By Vignatha Vinjam, Manjunath Kolavennu, Myna Vajha, Karthik Periyapattana Narayanaprasad
arXiv:2604.14908v2 Announce Type: replace-cross
Abstract: We study downlink beam and rate adaptation in a multi-user mmWave MISO system where multiple base stations (BSs), each using analog beamformi...
By Emre \"Ozy{\i}ld{\i}r{\i}m, Bar{\i}\c{s} Yayc{\i}, Umut Eren Akturk, Cem Tekin
arXiv:2607. 15404v1 Announce Type: cross Abstract: Interleaving mitigates burst errors but introduces decoding delay and removes temporal error structure that a channel-aware decoder could exploit.
By Bhaskar Krishnamachari
arXiv:2606. 00913v1 Announce Type: cross Abstract: Multi-arm bandit algorithms are increasingly used in online platforms, clinical trials, and social science experiments, but valid statistical inference on their performance remains an open challenge.
By Samya Praharaj, Chih-Yu Chang, Koulik Khamaru, Kelly W. Zhang
arXiv:2409. 18909v2 Announce Type: replace Abstract: Motivated by real-world applications that necessitate responsible experimentation, we introduce the problem of best arm identification (BAI) with minimal regret.
By Junwen Yang, Vincent Y. F. Tan, Tianyuan Jin
arXiv:2607. 09015v1 Announce Type: cross Abstract: We study contextual bandit problems with correlated arms and access to surrogate reward signals produced by a machine learning model, motivated by applications such as large language model (LLM) routing.
By Ajay Narayanan Sridhar, Ronak Singh, Mehrdad Mahdavi, Vijaykrishnan Narayanan
This paper studies Restless Multi‑Armed Bandits with individual penalty constraints for dynamic wireless networks, allowing each arm to have distinct performance limits such as energy, activation, or age of information. It introduces the Penalty‑Optimal Whittle (POW) index, which depends only on an arm’s transition kernel and its constraints, making it computable offline and independent of system‑wide parameters. The authors prove the POW index policy is asymptotically optimal, present a deep reinforcement learning method to learn the index online, and show through simulations that it outperforms existing policies.
By Nida Zamir, I-Hong Hou
arXiv:2609.06514v1 Announce Type: new
Abstract: Adaptive frequency hopping against predictive jamming must address both model uncertainty and policy exposure: the context-loss relationship may vary a...
By Yanbo Chen, Xinjing Zhou
arXiv:2608. 15660v1 Announce Type: cross Abstract: Federated learning (FL) enables privacy-preserving distributed model training but faces challenges from heterogeneous model architectures and limited communication resources at the network edge.
By Chenwang Liu, Yijun Liu, Chang Liu, Xu Zhang, Pengchao Han
arXiv:2609.26213v1 Announce Type: cross
Abstract: We study the multiplayer multi-armed bandit problem with information asymmetry under Bernoulli rewards, for three information structures: asymmetry i...
By Khang Nguyen, Ricardo Parada, William Chang
Exhaustively evaluating every candidate LLM configuration on every benchmark item to identify a high-performing one is costly. We formulate configuration selection as a cost-aware Bayesian bandit prob...
arXiv:2502. 13467v2 Announce Type: replace Abstract: The $K$-Max combinatorial multi-armed bandit problem arises in applications such as recommendation and distributed decision making, where the reward is determined by the maximum outcome among $K$ selected arms.
By Yu Chen, Siwei Wang, Longbo Huang, Wei Chen