arXiv AI By Baptiste Bonin, Maxime Heuillet, Audrey Durand

Offline A/B Testing of Slate Recommendation Systems with LLMs: Reducing the Dependency on Pre-Collected User Interaction Data

Read the original on arXiv AI →

The paper explores using large language models (LLMs) to generate pairwise preferences between slates for synthetic A/B testing of slate recommendation systems. It introduces a validation protocol that checks how well these synthetic preferences align with traditional RecSys metrics and satisfy preference axioms, and examines how LLM pre‑training and configuration influence preference articulation. By combining the synthetic preferences with a generalized Rao‑Kupper model, the authors show that LLM‑based A/B testing can recover stable ranking orderings across different utility weightings, offering a cost‑effective screening step before conducting live experiments.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 10

HyperTrace: Hypothesis-Based Preference Tracing for Online LLM Personalization

HyperTrace is a training‑free framework that personalizes large language models by tracing latent user preferences online. It maintains interpretable natural‑language hypotheses about short‑term intent and long‑term preferences, updating them with an SMC‑style reweighting process driven by an LLM‑based surrogate choice model. Experiments on PRISM and PersonaMem‑v2 demonstrate that HyperTrace improves response alignment, preference prediction, and profile consistency compared to strong online baselines.

By Jianzhi Shen, Keyu Mao, Minghao Shao, Chuanyang Jin, Yusong Wang, Ailiang Lin, Kotaro Funakoshi, Manabu Okumura, Tianmin Shu, Muhammad Shafique