The paper investigates online fair allocation of sequential items to agents with heterogeneous preferences, aiming to maximize generalized-mean welfare. In an i.i.d. arrival setting, a pure greedy algorithm achieves near-optimal “~O(1/T)” average regret without needing distributional knowledge. For nonstationary arrivals, the authors show that a single historical sample per distribution suffices to recover the same regret rate, using re-solving algorithms that remain robust to distribution shifts.
By Zongjun Yang, Rachitesh Kumar, Christian Kroer
arXiv:2607. 23310v1 Announce Type: cross Abstract: We study an online variant of discrete fair division under generalized assignment budget constraints.
By Saar Cohen, Nicholas Teh, Paul W. Goldberg, Michael J. Wooldridge
arXiv:2602.04125v2 Announce Type: replace-cross
Abstract: Modern digital platforms use contextual bandits to allocate valuable exposure and opportunities among competing participants. Fair treatment...
By Qingwen Zhang, Wenjia Wang
arXiv:2605. 01961v2 Announce Type: replace Abstract: Learning from human preference data is becoming a useful tool, from fine-tuning large language models to training reinforcement learning agents.
By Maheed H. Ahmed, Mahsa Ghasemi
arXiv:2601. 07144v3 Announce Type: replace-cross Abstract: Ensuring fairness in matching algorithms is a key challenge in allocating scarce resources and positions.
By Linus Bleistein, Mathieu Dagr\'eou, Francisco Andrade, Thomas Boudou, Aur\'elien Bellet
arXiv:2607. 23772v1 Announce Type: cross Abstract: We study a restless multi-armed bandit (RMAB) problem for a stochastic deadline scheduling application.
By Shakti Sharma, Rahul Meshram