Restless Bandits with Individual Penalty Constraints: Near-Optimal Indices and Deep Reinforcement Learning
Read the original on arXiv Machine Learning →This paper studies Restless Multi‑Armed Bandits with individual penalty constraints for dynamic wireless networks, allowing each arm to have distinct performance limits such as energy, activation, or age of information. It introduces the Penalty‑Optimal Whittle (POW) index, which depends only on an arm’s transition kernel and its constraints, making it computable offline and independent of system‑wide parameters. The authors prove the POW index policy is asymptotically optimal, present a deep reinforcement learning method to learn the index online, and show through simulations that it outperforms existing policies.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.