PersonaMem-v3 is a benchmark and evaluation harness designed to assess omni-platform personal intelligence for AI agents. It is built from over one million anonymized real-world engagement histories, covering social media, chatbots, calendars, and AI companions, and tracks user preferences and habits over time. The benchmark tests agents on personalization, LLM-powered recommendation, proactiveness, agentic tool use, and geo-temporal reasoning, evaluating their ability to infer holistic user understanding, personalize responses, rerank recommendations, follow user steering, and avoid inappropriate personalization.
By Bowen Jiang, Yuan Yuan, Zhuoqun Hao, Yuchen Liu, Maohao Shen, Sihao Chen, Gregory Wornell, Chris Callison-Burch, Lyle Ungar, Dan Roth, Qi Guo, Xiangjun Fan, Camillo J. Taylor, Hanchao Yu
arXiv:2607. 19739v1 Announce Type: cross Abstract: Although large language models (LLMs) have recently gained traction in recommender systems due to their strong reasoning capabilities and extensive world knowledge, previous LLM-based agents suffer from hallucination and context-length limitations, and thus are not suitable for full-ranking recommendation tasks.
By Mingdai Yang, Zhiwei Liu, Weizhi Zhang, Yibo Wang, Hao Peng, Philip Yu
arXiv:2607. 09988v1 Announce Type: cross Abstract: Recommendation systems, from traditional multi-stage to recent unified generative architectures, face challenges in incorporating diverse contextual signals, such as trending topics, breaking news, cultural events, and cross-surface user activities, into their ranking pipelines.
By Lei Shi, Di Wang, Harry Tran, Helsing Xu, Yuchen Lu, Dhara Ghodasara, Wilson Chaney, Xueting Liao, Jerry Yu, Huayu Ding, Mingze Gao, Shike Mei, Shuo Tang, Zhe Zhang, Jianming He, Abhishek Kumar, Haotian Wu, Hamed Firooz, Li Li
The rapid integration of large language model-based agents into recommender systems has driven a shift from static, ranking-based pipelines toward autonomous and interactive systems that can reason, plan, and act. This survey provides a comprehensive overview of this emerging landscape by introducing a unified taxonomy grounded in the level of autonomy and three core paradigms of agentic recommender systems: agent-assisted recommendation, agent-as-recommender, and agent-as-user-simulator.
Recommendation systems, from traditional multi-stage to recent unified generative architectures, face challenges in incorporating diverse contextual signals, such as trending topics, breaking news, cultural events, and cross-surface user activities, into their ranking pipelines. These systems are designed to consume structured behavioral signals with consistent schemas, and lack the reasoning capability to naturally process unstructured or heterogeneously formatted contextual information.
The paper introduces SARA, an industrial framework that scales articulated user rationales (AURs) for recommendation systems. It curates a high‑quality AUR dataset from 240 M users, trains a 7B‑parameter MLLM (SARA‑7B) to generate rationales for millions of authors, and integrates these generated rationales into a production ranking model (SARA‑Ranker). Offline and online experiments demonstrate that the system produces more specific, polarity‑consistent rationales and improves user engagement while reducing negative feedback.
By Haoke Xiao, Yueyang Liu, Yuhui Zhang, Xiang Chen, Yufei Liu, Jia Xu, Yalong Guan, Xiaolan Zhu, Xiaoyu Zhang, Shijun Wang, Shuang Yang, Zijie Meng, Zejian Zhang, Ruochen Yang, Xiangyu Wu, Tingting Gao, Han Li, Lantao Hu, Cheng Luo, Kun Gai
arXiv:2607. 25420v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used in recommender systems, but it is often unclear how much performance can be obtained from strong pre-trained backbones alone when they are placed inside a structured recommendation pipeline.
By Jiahao Tian, Zhenkai Wang
Although large language models (LLMs) have recently gained traction in recommender systems due to their strong reasoning capabilities and extensive world knowledge, previous LLM-based agents suffer from hallucination and context-length limitations, and thus are not suitable for full-ranking recommendation tasks. To circumvent these limitations through architectural design rather than modifying the LLM itself, we propose an agent-based recommendation framework, memory-based $\textbf{P}$ersonalized $\textbf{R}$ecommendation $\textbf{T}$ool learning via autonomous language $\textbf{A}$gents (PRTA), in which an LLM acts as a central planner interacting with multiple recommendation models as tools.
LIGE‑GR is a framework that transitions traditional ranking‑based recommender systems to a generative, listwise approach inspired by large language models. It extends existing pointwise recommendation models into a listwise generation system, enabling sequence‑level optimization without overhauling the entire infrastructure. Experiments on Instagram Reels and Facebook Video show modest gains in user time spent—1.14 % and 0.72 % respectively—while adding only slight inference overhead.
By Venkat Srinivas, Chenzhang He, Sam Woodmansee, Shawn Lian, Wenjie Hu, Renjie Jiang, Ziheng Huang, Xinyuan Zhang, Zhihao Zheng, Zhuoran Yu, Rui Li, Lei Yuan, Ziwei Li, Jimmy Jia, Mert Terzihan, Ekrem Kocaguneli, Yiming Liao, Zhichen Zhao, Yue Yin, Yue Weng, Wanlin Ma, Xufeng Cai, Weimiao Wu, Yezhou Huang, Du Zhang, Yukun Ding, Aaron Johnston, Yueming Wang, Zhaojie Gong, Yuting Zhang, Serena Li, Adithya Ganesh, Boying Liu, Haichuan Yang, Xialu Li, Matt Ma, Qunshu Zhang, John Joshua Miller, Praveen Rathinavelu, Cheng Huang, Aadhar Sachdeva, Josh Karns, Andres Aaron Gutierrez, Neil Agarwal, Gustas Pladis, Vladimir Batygin, Gopal Ray, Aditya Priyadarshi, Shantanu Patil, Zhe Wang, Penny Pan, Yiping Han, Arun Singh, Guangdeng Liao, Bi Xue, Xinyao Hu, Yang Song, Yisong Song, Meihong Wang, Haotian Wu, Deepak Agarwal, Ji Liu
The paper presents a pipeline for generating multi‑turn synthetic conversations and a self‑improvement loop that uses variance‑based contrastive optimization and a coding agent to refine planning and tool‑use in conversational recommendation agents. This approach improves agent quality by 8% over a manually optimized prompt and has been deployed at Spotify, where it accelerated development cycles. In production, the system achieved a 14% increase in user listening, a 5% rise in weekly active users, and a 5% reduction in skip rate compared to a prior session‑only experience.
By Enrico Palumbo, Alexandre Tamborrino, Victor Ode, Ben Lacker, Adri\`a Casas Escoda, Jeremy Hopple, Marcus Better, James Leoni, Hugo Galv\~ao, Hugues Bouchard, Mounia Lalmas, Jos\'e Luis Redondo Garc\'ia, Abenezer Abebe, Ann Clifton, Anton Blomberg, Henrik Lindstr\"om, Dani Doro, Christine Doig Cardet
ReMem is a new recommendation agent framework that rethinks perception and memory for long-context recommendation tasks. It replaces raw HTML parsing with OCR-based multimodal perception from screenshots, extracting structured information in a platform-agnostic way. The framework also introduces a chunk-wise sequential memory update strategy and a multi-memory GRPO variant to efficiently model evolving user preferences over arbitrarily long interaction histories, achieving a 5.16% average improvement over state-of-the-art baselines on three recommendation agent tasks.
By Haohao Qu, Yongcheng Jing, Chun Hin Chan, Shanru Lin, Wenqi Fan, Dacheng Tao
arXiv:2603. 21613v2 Announce Type: replace-cross Abstract: Recommender agents built on Large Language Models offer a promising paradigm for personalized recommendation.
By Tianyi Li, Zixuan Wang, Guidong Lei, Xiaodong Li, Hui Li