arXiv:2606. 06178v1 Announce Type: new Abstract: Large language models (LLMs) present a trade-off between performance and cost, where more powerful models incur greater expense.
By Jiahao Zeng, Ming Tang, Ningning Ding
arXiv:2609.37800v1 Announce Type: cross
Abstract: Many recommender services repeatedly encounter cold-start cohorts, where new users arrive with little or no interaction history. This creates two cha...
By Serafima Lebedeva, Sumantrak Mukherjee, Ali Arshad Sadal, Ilias Ek\c{s}i, Rahul Sharma, Julia Mueller, Theresa Dombrowski, Jakob Karolus, Viktor Bengs, Eyke H\"ullermeier, Sebastian Vollmer
arXiv:2609.38860v1 Announce Type: cross
Abstract: Learning from human preferences is central to large language model (LLM) alignment, but human preference annotation is costly. Active preference lear...
By Zhongman Du, Huiming Zhang, Haodong Zhu, Baochang Zhang
The paper introduces a new variant of the multi‑armed bandit problem that incorporates dueling feedback—pairwise comparisons of model responses—and heterogeneous sampling costs to identify the best large language model (LLM) from a set with varying query costs. Assuming a Condorcet winner, the authors propose a Track‑and‑Stop style algorithm that guarantees asymptotically optimal cost as the error probability approaches zero. Extensive experiments on synthetic and real‑world data show that this cost‑aware approach consistently outperforms both classical cost‑unaware algorithms and other cost‑aware extensions.
By Sarvesh Gharat, Nikhil Karamchandani, Jayakrishnan Nair
arXiv:2608. 06559v1 Announce Type: new Abstract: Contextual bandits offer a natural framework for sample-efficient personalization, but practical deployment remains difficult under sparse, biased interaction data, unreliable uncertainty estimates, and severe cold starts.
By Devansh Gupta, Shiv Tavker, Dmitry Efimov, Suchitra Sathyanarayana, Gitanjali Bhutani, Boris N. Oreshkin
The paper introduces Latent Order Bandits (LOB), a new bandit framework that relaxes the strict assumptions of traditional latent bandits by only requiring a partial order of action preferences within each latent state. LOB allows instances sharing the same state to have different reward distributions as long as the action ranking remains consistent, making it suitable for scenarios like user groups on streaming services who agree on genre preferences but rate differently. The authors present an upper‑confidence bound algorithm for both total and partial latent orders, provide regret bounds, and propose a posterior‑sampling variant that empirically outperforms full‑prior latent bandits when reward scales vary across instances sharing the same latent state.
By Emil Carlsson, Newton Mwai, Fredrik D. Johansson