We study repeated bidding in multi-unit discriminatory (pay-as-bid) auctions for a single bidder with per-round utility equal to value minus $α$ times payment, where $α\in[0,1]$ is a cost-of-capital parameter. The bidder aims to maximize cumulative utility over $T$ rounds subject to a total budget $B$.
We study the contextual combinatorial semi-bandit (CCSB) problem with general reward function approximation. At each round, the learner observes a context, selects a combinatorial action consisting of a subset of basic arms, and receives the reward of each selected arm; the goal is to maximize the cumulative reward over time.
arXiv:2606. 29252v1 Announce Type: new Abstract: We study repeated bidding in multi-unit discriminatory (pay-as-bid) auctions for a single bidder with per-round utility equal to value minus $\alpha$ times payment, where $\alpha\in[0,1]$ is a cost-of-capital parameter.
By Negin Golrezaei, Sourav Sahoo
arXiv:2607. 13686v1 Announce Type: new Abstract: We study the contextual combinatorial semi-bandit (CCSB) problem with general reward function approximation.
By Hao Qin, Chicheng Zhang
arXiv:2608. 11560v1 Announce Type: new Abstract: Personalizing marketing messages with contextual multi-armed bandits (CMABs) drives real business value, yet the objective that ultimately matters - a downstream conversion - is observed only weeks later, too late to drive online learning.
By Sang Su Lee, Vineeth Loganathan, Shishir Dash, Vijay Raghavan
arXiv:2506. 03802v2 Announce Type: replace Abstract: We introduce a learning problem in a generalized two-sided matching market, where agents select actions to interact with their match.
By Andreas Athanasopoulos, Christos Dimitrakakis
The paper investigates preference-based bandits where a learner selects pairs of arms and receives binary preference feedback modeled by Bradley–Terry. It introduces the locally sensitive eluder dimension, a new complexity measure for logistic preference feedback, and proposes the GINOP algorithm that uses log-loss confidence sets to balance optimism and exploration. The authors prove a first-order regret bound showing that learning with preference feedback can be as statistically efficient as learning from direct rewards, and they validate their theory with empirical experiments.
By Ahmed Ben Yahmed (CREST, ENSAE Paris, FAIRPLAY), Marc Abeille (FAIRPLAY), Cl\'ement Calauz\`enes (FAIRPLAY)
Personalizing marketing messages with contextual multi-armed bandits (CMABs) drives real business value, yet the objective that ultimately matters - a downstream conversion - is observed only weeks later, too late to drive online learning. Teams therefore train the bandit on a fast proxy reward, and separately must judge whether a contextual bandit is worth its complexity over sending one best message.
arXiv:2602. 09456v2 Announce Type: replace Abstract: We propose an algorithmic framework, Offline Estimation to Decisions (OE2D), that efficiently reduces contextual bandit learning with general reward function approximation to offline regression.
By Hao Qin, Chicheng Zhang
arXiv:2609.26213v1 Announce Type: cross
Abstract: We study the multiplayer multi-armed bandit problem with information asymmetry under Bernoulli rewards, for three information structures: asymmetry i...
By Khang Nguyen, Ricardo Parada, William Chang
arXiv:2602.04125v2 Announce Type: replace-cross
Abstract: Modern digital platforms use contextual bandits to allocate valuable exposure and opportunities among competing participants. Fair treatment...
By Qingwen Zhang, Wenjia Wang
arXiv:2602.10727v3 Announce Type: replace
Abstract: Rising Multi-Armed Bandits (RMABs) model sequential decision problems where each arm's expected reward improves with repeated pulls. In such proble...
By Seockbean Song, Chenyu Gan, Youngsik Yoon, Siwei Wang, Wei Chen, Jungseul Ok