arXiv AI

PromptPack: Scaling LLM Annotation Agents for Online Recommendation

arXiv:2607. 20528v1 Announce Type: new Abstract: Online recommendation platforms increasingly use Large Language Models (LLMs) to extract structured features from ad creatives.

arXiv AI
6d ago

Bootstrapping Conversational Recommendation Agents At Spotify: Synthetic Data Generation and Self-Improvement Loops

The paper presents a pipeline for generating multi‑turn synthetic conversations and a self‑improvement loop that uses variance‑based contrastive optimization and a coding agent to refine planning and tool‑use in conversational recommendation agents. This approach improves agent quality by 8% over a manually optimized prompt and has been deployed at Spotify, where it accelerated development cycles. In production, the system achieved a 14% increase in user listening, a 5% rise in weekly active users, and a 5% reduction in skip rate compared to a prior session‑only experience.

By Enrico Palumbo, Alexandre Tamborrino, Victor Ode, Ben Lacker, Adri\`a Casas Escoda, Jeremy Hopple, Marcus Better, James Leoni, Hugo Galv\~ao, Hugues Bouchard, Mounia Lalmas, Jos\'e Luis Redondo Garc\'ia, Abenezer Abebe, Ann Clifton, Anton Blomberg, Henrik Lindstr\"om, Dani Doro, Christine Doig Cardet
arXiv Machine Learning
1d ago

Which LLM to pick? Online Active Model Selection for Large Language Models

The paper introduces ONLINE LLM PICKER, a framework for active model selection of large language models in streaming settings. It selects the most informative prompts for annotation within a limited budget, enabling the identification of the best or near‑best model among many candidates. Experiments on 10 datasets and over 130 language models show up to 71.67% savings in annotation cost and a reduction in regret by up to 2.51× when using the chosen model for sequential generation.

By Alessandro Turrin, Patrik Okanovic, Torsten Hoefler, Nezihe Merve G\"urel
Hugging Face Trending Papers
Jul 22

Personalized Recommendation Tool Learning via Autonomous Language Agents

Although large language models (LLMs) have recently gained traction in recommender systems due to their strong reasoning capabilities and extensive world knowledge, previous LLM-based agents suffer from hallucination and context-length limitations, and thus are not suitable for full-ranking recommendation tasks. To circumvent these limitations through architectural design rather than modifying the LLM itself, we propose an agent-based recommendation framework, memory-based $\textbf{P}$ersonalized $\textbf{R}$ecommendation $\textbf{T}$ool learning via autonomous language $\textbf{A}$gents (PRTA), in which an LLM acts as a central planner interacting with multiple recommendation models as tools.

arXiv AI
Jul 23

Personalized Recommendation Tool Learning via Autonomous Language Agents

arXiv:2607. 19739v1 Announce Type: cross Abstract: Although large language models (LLMs) have recently gained traction in recommender systems due to their strong reasoning capabilities and extensive world knowledge, previous LLM-based agents suffer from hallucination and context-length limitations, and thus are not suitable for full-ranking recommendation tasks.

By Mingdai Yang, Zhiwei Liu, Weizhi Zhang, Yibo Wang, Hao Peng, Philip Yu
arXiv AI
Aug 10

Shape Your Feed: An LLM-based Agentic System for Conversational Recommendation

arXiv:2608. 06632v1 Announce Type: new Abstract: Industrial recommendation systems predominantly adopt a passive ranking paradigm that infers user preferences from implicit behavioral signals (e.

By Ziyun Xu, Bosen Ding, Yue Zhang, Ji Qi, Qingyuan Song, Jizhou Huang, Liwei Wang, Jefferey Santelli, Yue Weng, Qichao Que, Zhenheng Yang, Junfeng Pan, Linhong Zhu
arXiv Machine Learning
Aug 28

A Survey of LLM Prompt Datasets: Taxonomy, Linguistic Patterns, and Practical Uses

The paper presents a survey of 129 public large language model (LLM) prompt datasets, totaling over 1.22 TB and 673 million instances, and introduces a unified taxonomy for them. By analyzing seven datasets in depth, the authors identify lexical, syntactic, and semantic patterns that differentiate prompts from general text, and evaluate these patterns for tasks such as prompt filtering, source domain routing, and response quality assessment. They demonstrate that a 63‑dimensional linguistic feature set extracted on a CPU can match over 91 % of the F1 score of GPU‑based sentence embeddings while halving latency, and that structural features can effectively route prompts across datasets, though they may negatively impact response quality when prompt length is controlled.

By Yuanming Zhang, Yan Lin, Arijit Khan, Huaiyu Wan
arXiv Machine Learning
Jun 25

TokenMinds: Pretrained User Tokens and Embeddings for User Understanding in Large Recommender Systems

arXiv:2606. 25147v1 Announce Type: cross Abstract: User modeling in industrial recommender systems typically produces dense embeddings, which suffer from representational constraints inherent to fixed-dimensional vectors.

By Qingyun Liu, Bo Yan, Yang Liu, Yuji Roh, Ekansh Sharma, Likang Yin, Emma Olowo, Min-hsuan Tsai, Yuxuan Li, Diego Uribe, Saksham Aggarwal, Siqi Wu, Yuan Hao, Vikas Kedigehalli, Lukasz Heldt, Lichan Hong, Li Wei, Xinyang Yi