Impatient Bandits: Optimizing for the Long-Term Without Delay
arXiv:2501. 07761v2 Announce Type: replace-cross Abstract: Increasingly, recommender systems are tasked with improving users' long-term satisfaction.
arXiv:2501. 07761v2 Announce Type: replace-cross Abstract: Increasingly, recommender systems are tasked with improving users' long-term satisfaction.
arXiv:2608. 03382v1 Announce Type: cross Abstract: Multi-armed bandit algorithms, especially Thompson sampling, are widely used in online recommendation.
arXiv:2609.00251v1 Announce Type: new Abstract: As people increasingly interact with LLM assistants in daily life, continually adapting to individual preferences has become essential for effective lo...
arXiv:2608. 06559v1 Announce Type: new Abstract: Contextual bandits offer a natural framework for sample-efficient personalization, but practical deployment remains difficult under sparse, biased interaction data, unreliable uncertainty estimates, and severe cold starts.
arXiv:2609.15598v1 Announce Type: cross Abstract: Generative recommendation has emerged as a promising end-to-end paradigm for personalized recommendation. However, user preferences continuously evol...
arXiv:2511. 22130v2 Announce Type: replace Abstract: To navigate ever-shifting real-world environments, agents must grapple with incomplete knowledge and adapt their strategies through experience.
arXiv:2607. 14192v1 Announce Type: new Abstract: As recommender systems mature in the past few years, their optimization objectives have evolved from a primary focusing on short-term behavioral signals to a broader emphasis on long-term user engagement and retention.
The paper introduces Latent Order Bandits (LOB), a new bandit framework that relaxes the strict assumptions of traditional latent bandits by only requiring a partial order of action preferences within each latent state. LOB allows instances sharing the same state to have different reward distributions as long as the action ranking remains consistent, making it suitable for scenarios like user groups on streaming services who agree on genre preferences but rate differently. The authors present an upper‑confidence bound algorithm for both total and partial latent orders, provide regret bounds, and propose a posterior‑sampling variant that empirically outperforms full‑prior latent bandits when reward scales vary across instances sharing the same latent state.
arXiv:2606. 17276v1 Announce Type: cross Abstract: Generative recommendation (GR) has emerged as a promising direction for recommender systems.
arXiv:2606. 00846v1 Announce Type: new Abstract: Users increasingly face the challenge of selecting an appropriate LLM for a given task from a rapidly growing pool of LLMs, each with distinct but often opaque latent properties.
arXiv:2606. 23933v1 Announce Type: cross Abstract: We study non-stationary linear contextual bandits where the reward model drifts over time, rendering classical contextual bandit algorithms brittle because historical data becomes systematically biased.
arXiv:2601. 09974v2 Announce Type: replace Abstract: Personalizing Large Language Models typically relies on static retrieval or one-time adaptation, assuming user preferences remain invariant over time.