arXiv Machine Learning By Dingsu Wang, Filip Ryzner, Kelly He, Armando Ordorica, David Woo, Aditya Mantha, Liyao Lu, Usha Amrutha Nookala, Haoran Guo, Jiacong He, Olafur Gudmundsson, Matt Chun, Krystal Benitez, Dhruvil Deven Badani, Yijie Dylan Wang

Long-term User Engagement Optimization through Model-agnostic Downstream Rewards Learning

Read the original on arXiv Machine Learning →

arXiv:2607. 14192v1 Announce Type: new Abstract: As recommender systems mature in the past few years, their optimization objectives have evolved from a primary focusing on short-term behavioral signals to a broader emphasis on long-term user engagement and retention.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 24

COPE: Continual Personalization of LLMs under Sparse User Feedback via User Embeddings and Self-Evaluation

COPE (Continual Optimization with Personalized embedding and self-Evaluation) is a new framework that continually personalizes large language models using learnable user embeddings and self‑evaluation to generate proxy rewards. It integrates preference capture, self‑evaluation calibration, and personalized response optimization into a single update step, allowing continuous model updates even when explicit user feedback is sparse. Experiments demonstrate that COPE outperforms both training‑free and training‑based baselines, remains complementary to Retrieval‑Augmented Prompting, and shows reliable self‑evaluation, meaningful preference patterns, stable general capabilities, and robustness to shifting preferences and alternative evaluators.

By Ruike Cao, Fugen Yao, Liang Dong, Jian Xu, Guanjun Jiang, Li Xiao
arXiv AI
4d ago

Challenges and Solutions for Bandits in the Wild: Warm-Started Mixture Bandits for Cross-Cohort Slate Recommendation

arXiv:2609.37800v1 Announce Type: cross Abstract: Many recommender services repeatedly encounter cold-start cohorts, where new users arrive with little or no interaction history. This creates two cha...

By Serafima Lebedeva, Sumantrak Mukherjee, Ali Arshad Sadal, Ilias Ek\c{s}i, Rahul Sharma, Julia Mueller, Theresa Dombrowski, Jakob Karolus, Viktor Bengs, Eyke H\"ullermeier, Sebastian Vollmer
arXiv Machine Learning
Jul 31

Personalized RewardBench: Evaluating Reward Models with Human Aligned Personalization

arXiv:2604. 07343v2 Announce Type: replace-cross Abstract: Pluralistic alignment has emerged as a critical frontier in the development of Large Language Models (LLMs), with reward models (RMs) serving as a central mechanism for capturing diverse human values.

By Qiyao Ma, Dechen Gao, Rui Cai, Boqi Zhao, Hanchu Zhou, Junshan Zhang, Zhe Zhao
arXiv AI
Sep 10

SequenceO1: End-to-End Ultra-Long (100K) Sequence Modeling in Recommendation with Low-Rank Caching

SequenceO1 is an end‑to‑end framework that enables ultra‑long (up to 100K interactions) sequence modeling for recommendation systems. It compresses raw user histories into a fixed‑size sketch using Sketch Attention and then models short‑term and long‑term interests with Target‑to‑History Cross Attention. The system incorporates low‑rank caching, batching, pipeline lift, and a FlashSA kernel to keep training and inference efficient, achieving consistent offline and online performance gains when deployed at full traffic on Douyin.

By Lin Guan, Jia-Qi Yang, Zhishan Zhao, Jiaqi Huang, Hangyu Wang, Longbin Li, Beichuan Zhang, Haonan Jiang, Jinan Ni, Xiangyu Fan, Xiaowen Li, Ziyao Ren, Yuhang Qi, Xiaolong Zhu, Xuanyuan Luo, Qiwei Chen, Yi Cheng, Lele Yu