GISTBench: Evaluating LLM User Understanding via Evidence-Based Interest Verification
Read the original on arXiv AI →GISTBench is a benchmark designed to assess how well Large Language Models understand users by extracting and verifying user interests from their interaction histories in recommendation systems. It introduces two new metric families—Interest Groundedness (IG) and Interest Specificity (IS)—to measure the accuracy and distinctiveness of LLM-generated user profiles. The benchmark includes a synthetic dataset built from real user interactions on a global short‑form video platform, validated through user surveys, and evaluates a range of open‑weight and proprietary LLMs, uncovering limitations in their ability to count and attribute engagement signals.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.