$\Psi$-Bench: Evaluating Persona-Sensitive Influencing in Persuasive Dialogues
arXiv:2606. 02754v1 Announce Type: new Abstract: Personalization is a crucial capability of modern language agents.
arXiv:2607. 20471v1 Announce Type: new Abstract: Personalization, the act of varying a message to induce action from a specific receiver while keeping sender, channel, and time fixed, has a long tradition in psychology and marketing as a two-party problem in which sender and receiver have independent objectives.
arXiv:2606. 02754v1 Announce Type: new Abstract: Personalization is a crucial capability of modern language agents.
arXiv:2604. 07343v2 Announce Type: replace-cross Abstract: Pluralistic alignment has emerged as a critical frontier in the development of Large Language Models (LLMs), with reward models (RMs) serving as a central mechanism for capturing diverse human values.
arXiv:2606. 06614v1 Announce Type: cross Abstract: Despite growing interest, most evaluations of large language models' (LLMs') personalization abilities have relied on synthetic data.
The paper introduces SalesLLM, a bilingual (Chinese/English) benchmark for evaluating large language models (LLMs) in realistic sales dialogues. It comprises 30,074 scripted configurations and 1,805 curated multi‑turn scenarios from Financial Services and Consumer Goods, with controllable difficulty and personas. An automatic evaluation pipeline uses an LLM judge for sales‑process progress and fine‑tuned BERT classifiers for end‑of‑dialogue buying intent, while a user model, CustomerLM, is trained to improve simulation fidelity. SalesLLM scores correlate strongly with human ratings (Pearson r = 0.86) and reveal that top Chinese LLMs match junior‑to‑intermediate human salespeople but not experts, with cross‑lingual consistency remaining poor.
arXiv:2606. 05330v1 Announce Type: cross Abstract: Large language models can shift human beliefs across high-stakes domains, but most persuasion studies rely on pre/post belief change.
arXiv:2608.28833v1 Announce Type: new Abstract: While Large language models (LLMs) incorporate user personalization signals to improve usability and helpfulness, they increasingly shift from providin...
The paper presents a method for fine‑tuning a large language model (LLM) recommender to generate personalized, non‑harmful explanations for its recommendations. By training two LLM‑judge reward models and using constrained GRPO, the authors achieve a significant increase in the PASS rate for all three criteria, from 0.649 to 0.956 on their own judges and from 0.677 to 0.931 on an independent judge. The fine‑tuned model maintains its original recommendation performance, demonstrating that LLM‑based recommenders can be adapted to complex tasks without loss of effectiveness.
arXiv:2606. 21097v2 Announce Type: replace-cross Abstract: Deploying highly capable personalized conversational agents in resource-constrained or privacy-sensitive environments remains a significant challenge.
arXiv:2608.30023v1 Announce Type: cross Abstract: Generative engines such as ChatGPT, Gemini, and Perplexity answer buyer questions directly and name a shortlist of brands inside the answer. Studying...
arXiv:2601.19435v2 Announce Type: replace-cross Abstract: Sustainable monetization of large language models (LLMs) remains a critical open challenge. Traditional search advertising, which relies on s...
The study evaluates whether large language models (LLMs) used as synthetic personas can predict real audience responses to marketing copy. Using thousands of headline A/B tests from the Upworthy Research Archive, the authors compare a ten-persona panel grounded in real audience demographics to a no-persona zero‑shot baseline that asks the model for a typical reader’s click likelihood. Results show that the no‑persona baseline outperforms the persona‑based approach, with higher predictive validity and top‑1 accuracy, indicating that forcing the model to role‑play specific personas introduces bias and noise.
The paper introduces a lifecycle framework for LLM-as-a-Judge systems used to evaluate recommendation explanations at Netflix. It outlines four phases—Birth, Training, Deployment, and Monitoring—detailing how each stage addresses specific technical and operational challenges. The authors report that after five weeks of A/B testing, judge-aligned explanations increased novel content viewing and successful browse-to-play sessions without quality takedowns.