arXiv AI

GISTBench: Evaluating LLM User Understanding via Evidence-Based Interest Verification

GISTBench is a benchmark designed to assess how well Large Language Models understand users by extracting and verifying user interests from their interaction histories in recommendation systems. It introduces two new metric families—Interest Groundedness (IG) and Interest Specificity (IS)—to measure the accuracy and distinctiveness of LLM-generated user profiles. The benchmark includes a synthetic dataset built from real user interactions on a global short‑form video platform, validated through user surveys, and evaluates a range of open‑weight and proprietary LLMs, uncovering limitations in their ability to count and attribute engagement signals.

arXiv Machine Learning
Jul 31

Personalized RewardBench: Evaluating Reward Models with Human Aligned Personalization

arXiv:2604. 07343v2 Announce Type: replace-cross Abstract: Pluralistic alignment has emerged as a critical frontier in the development of Large Language Models (LLMs), with reward models (RMs) serving as a central mechanism for capturing diverse human values.

By Qiyao Ma, Dechen Gao, Rui Cai, Boqi Zhao, Hanchu Zhou, Junshan Zhang, Zhe Zhao
arXiv Machine Learning
Jun 25

TokenMinds: Pretrained User Tokens and Embeddings for User Understanding in Large Recommender Systems

arXiv:2606. 25147v1 Announce Type: cross Abstract: User modeling in industrial recommender systems typically produces dense embeddings, which suffer from representational constraints inherent to fixed-dimensional vectors.

By Qingyun Liu, Bo Yan, Yang Liu, Yuji Roh, Ekansh Sharma, Likang Yin, Emma Olowo, Min-hsuan Tsai, Yuxuan Li, Diego Uribe, Saksham Aggarwal, Siqi Wu, Yuan Hao, Vikas Kedigehalli, Lukasz Heldt, Lichan Hong, Li Wei, Xinyang Yi