arXiv:2606. 27917v1 Announce Type: new Abstract: Contextual bandits with graph-structured arms arise in recommendation, citation retrieval, and social advertising, where arms connected on a graph tend to share reward signal.
By Joyanta Jyoti Mondal, Ibne Farabi Shihab, Anuj Sharma
arXiv:2604. 00531v2 Announce Type: replace Abstract: Multi-task representation learning exploits the shared structure among related tasks by learning a common latent representation, thereby improving sample efficiency.
By Jiabin Lin, Shana Moothedath
The paper studies high‑dimensional linear contextual bandits with knapsack constraints (CBwK), aiming to exploit sparsity for tighter regret bounds. It introduces an online hard‑thresholding estimator integrated into a primal‑dual framework, achieving sub‑linear regret that grows only logarithmically with the feature dimension. Under either a diverse‑covariate or margin condition, the regret improves to τ‑dependent rates, and when both hold simultaneously, a dual resolving scheme yields an even tighter bound. The approach also recovers optimal rates for high‑dimensional contextual bandits without knapsacks, and experiments demonstrate its practical effectiveness.
By Wanteng Ma, Dong Xia, Jiashuo Jiang
arXiv:2608. 04324v1 Announce Type: cross Abstract: This paper studies generalized low-rank matrix bandits with multiple prioritized objectives.
By Bo Xue, Ji Cheng, Haodong Jing, Hongzong Li, Shuang Qiu
arXiv:2606. 11968v1 Announce Type: new Abstract: This paper studies efficient online algorithms for multinomial logistic bandits (MLogB), where the feedback distribution over $K+1$ outcomes follows a multinomial logistic model of $d$-dimensional action vectors.
By Linzhe He, Yu-Jie Zhang, Sifan Yang, Lijun Zhang
arXiv:2606. 31449v1 Announce Type: new Abstract: We investigate the contextual slate bandit problem with generalized linear rewards under limited adaptivity.
By Tanmay Goyal, Sukruta Prakash Midigeshi, Gaurav Sinha