arXiv AI

Large Behavior Model: A Promptable Digital Twin of the Retail Customer

arXiv:2607. 06993v1 Announce Type: new Abstract: Customer behavior modeling underpins recommendation, marketing, and decision support, yet existing approaches either optimize predictive accuracy without explaining decisions or simulate users without grounding them in real behavioral data.

arXiv Machine Learning
Sep 17

Behavior2Value: Benchmarking and Empowering LLMs for Consumer Value Measurement from E-commerce Behaviors

The paper introduces the Behavior-to-Value (B2V) task, which seeks to identify consumer values from e-commerce behavioral trajectories. It presents the E-commerce Consumption Value Taxonomy (ECVT) and the B2V-Bench dataset, derived from anonymized Taobao logs and covering 25 purchase behaviors with associated value orientations. A new model, B2V-Verifier, is proposed to improve value measurement accuracy, achieving a 34% boost in multi-label classification over strong LLM baselines.

By Peixuan Hou, Bin Chen, Li He, Jian Xu, Bo Zheng, Xiuli Ma, Guojie Song
arXiv AI
Sep 25

Beyond Surface Style: Aligning Multi-Turn User Simulators with Behavioral Consistency

The paper introduces TRACER, a multi‑turn user simulator that models evolving user intent and aligns simulated behavior with real interaction trajectories. TRACER is trained first with supervised fine‑tuning on real dialogues and then with reinforcement learning that uses hierarchical outcome‑ and trajectory‑level rewards to address reward sparsity and credit assignment. In real customer‑service sessions, TRACER‑7B outperforms the best baseline by 11.4 conversion F1, achieves the lowest group‑level conversion‑rate error and semantic trajectory distance, and generalizes to out‑of‑distribution scenarios, while human Turing tests show its conversations appear natural. The authors also present the Dynamic Marketing Benchmark, which evaluates both persuasion effectiveness and response quality of large language models through simulated interactions, demonstrating that higher response quality does not always lead to higher conversion rates.

By Geng Chen, Ruotong Pan, Zhirui Yang, Qiqi He, Jiawei Chen, Zhang Yunfei, Chongyuan Chen, Minxuan Lv, Zheng Yang, Win-Bin Huang, Xiangyu Wu, Wenwu Ou
arXiv Computation and Language
Sep 10

See Better, Foresee Better, Act Wiser: Physically Grounded Proactive Modeling and Decision Making

arXiv:2606.03371v4 Announce Type: replace Abstract: Reliable proactive agents must choose an action and judge whether current evidence is sufficient to act. We study retail service from sparse third-...

By Honghui Zhang, Anna Min, Chenmeinian Guo, Yujia Zhang, Yichen Yu, Zezhou Zhang, Guanyu Liu, Yongming Qin, Chongguo Song, Mengyue Yang, Lei Yu, Tianyu Shi
arXiv AI
Sep 1

Beyond Ranking Accuracy: Evaluating LLM-Cited Feature Rationales for Next Basket Repurchase Recommendation

arXiv:2608.30333v1 Announce Type: cross Abstract: Next-basket repurchase recommendation is commonly formulated as a ranking task: given a customer's purchase history, the system ranks previously purc...

By Yanan Cao, Anay Dombe, Murali Mohana Krishna Dandu, Shreeranjani Srirangamsridharan, Sinduja Subramaniam, Yogananth Mahalingam, Evren Korpeoglu, Kannan Achan
arXiv AI
Aug 3

DRIP-R: A Benchmark for Decision-Making and Reasoning Under Real-World Policy Ambiguity in the Retail Domain

arXiv:2605. 07699v2 Announce Type: replace-cross Abstract: LLM-based agents are increasingly deployed for routine but consequential tasks in real-world domains, where their behavior is governed by inherently ambiguous domain policies that admit multiple valid interpretations.

By Hsuvas Borkakoty, Sebastian Pohl, Cheng Wang, Bei Chen, Yufang Hou