arXiv AI

Ottilie: A Socially Intelligent Virtual Host for Sales-Driven Live Commerce

arXiv Computation and Language
Sep 17

"If I Had to Buy Just ONE: Galaxy S26 Ultra": Auditing AI-Generated Product Recommendations

arXiv:2609.18729v1 Announce Type: cross Abstract: Consumers increasingly use AI chatbots for advice on what to buy. With companies like OpenAI and Google monetising their AI through advertising, this...

By Lucas G. Uberti-Bona Marin, Thales Bertaglia, Giovanni Astante, Bram Rijsbosch, Gijs van Dijck, Anik\'o Hann\'ak, Gerasimos Spanakis, Konrad Kollnig
arXiv AI
Aug 10

Shape Your Feed: An LLM-based Agentic System for Conversational Recommendation

arXiv:2608. 06632v1 Announce Type: new Abstract: Industrial recommendation systems predominantly adopt a passive ranking paradigm that infers user preferences from implicit behavioral signals (e.

By Ziyun Xu, Bosen Ding, Yue Zhang, Ji Qi, Qingyuan Song, Jizhou Huang, Liwei Wang, Jefferey Santelli, Yue Weng, Qichao Que, Zhenheng Yang, Junfeng Pan, Linhong Zhu
arXiv AI
Aug 20

Bridging Search and CRM: Productionizing AI Product Research Agents for Customer Re-Engagement

The paper introduces a production-ready framework that connects e‑commerce search and CRM systems via AI‑powered Product Research Agents. These agents detect users with exploratory purchase intent, perform multi‑agent research using behavioral data, external knowledge, and catalog information, and then send personalized product recommendations through WhatsApp. In a 23‑day deployment, the system sent about 15,000 notifications, achieving higher click‑through rates than standard campaigns and generating downstream purchases and GMV gains.

By Mandar Kulkarni, Pooja A., Samir Shah
arXiv AI
1d ago

RealWorldShop: Benchmarking and Improving Conversational Shopping Agents in Real-World E-commerce

RealWorldShop introduces a new benchmark for conversational shopping agents, featuring 3.28 million grounded products, structured shopping episodes, a profile‑grounded user simulator, and role‑play evaluation. Analysis reveals that existing systems generate locally plausible responses but struggle with state tracking, constraint updating, and grounded convergence, especially in ambiguous or multi‑intent scenarios. The authors propose REALSHOP_AGENT, a session‑control framework with explicit state management, shopping‑flow control, catalog‑grounded retrieval, and runtime guards, which consistently outperforms strong baselines on the benchmark.

By Xinwei Yang, Kelong Mao, Yudong Guo, Sulong Xu, Simiu Gu, Chen Huang, Wenqiang Lei
arXiv AI
Jul 15

Operationalising Multi-Dimensional Evaluation for Conversational Agents: A Scalable, Governed Pipeline with Selective Re-evaluation and Model Benchmarking

arXiv:2607. 12085v1 Announce Type: new Abstract: Evaluating retail conversational agents requires methods beyond lexical-overlap metrics to assess intent alignment, factuality, helpfulness, clarity, tone, and overall response quality.

By Niranjan Kumar M, Balaji Nagarajan, Karthik Nair, Faysal Satter, Nithin Surendran
arXiv AI
Aug 20

The Lifecycle of LLM-as-a-Judge for Large-Scale Recommendation Explanations

The paper introduces a lifecycle framework for LLM-as-a-Judge systems used to evaluate recommendation explanations at Netflix. It outlines four phases—Birth, Training, Deployment, and Monitoring—detailing how each stage addresses specific technical and operational challenges. The authors report that after five weeks of A/B testing, judge-aligned explanations increased novel content viewing and successful browse-to-play sessions without quality takedowns.

By Emma Yanyang Kong, JJ Tan, Ishan Gupta, Lars Olds, Claire Campbell, David Fagnan, Veli Balin, Rohan Gosain, Louis Garcia, Minsu Jang
arXiv AI
4d ago

Bootstrapping Conversational Recommendation Agents At Spotify: Synthetic Data Generation and Self-Improvement Loops

The paper presents a pipeline for generating multi‑turn synthetic conversations and a self‑improvement loop that uses variance‑based contrastive optimization and a coding agent to refine planning and tool‑use in conversational recommendation agents. This approach improves agent quality by 8% over a manually optimized prompt and has been deployed at Spotify, where it accelerated development cycles. In production, the system achieved a 14% increase in user listening, a 5% rise in weekly active users, and a 5% reduction in skip rate compared to a prior session‑only experience.

By Enrico Palumbo, Alexandre Tamborrino, Victor Ode, Ben Lacker, Adri\`a Casas Escoda, Jeremy Hopple, Marcus Better, James Leoni, Hugo Galv\~ao, Hugues Bouchard, Mounia Lalmas, Jos\'e Luis Redondo Garc\'ia, Abenezer Abebe, Ann Clifton, Anton Blomberg, Henrik Lindstr\"om, Dani Doro, Christine Doig Cardet