arXiv AI By Wachiravit Modecrua, Krittin Pachtrachai, Touchapon Kraisingkorn

Large Behavior Model: A Promptable Digital Twin of the Retail Customer

Read the original on arXiv AI →

arXiv:2607. 06993v1 Announce Type: new Abstract: Customer behavior modeling underpins recommendation, marketing, and decision support, yet existing approaches either optimize predictive accuracy without explaining decisions or simulate users without grounding them in real behavioral data.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 17

Behavior2Value: Benchmarking and Empowering LLMs for Consumer Value Measurement from E-commerce Behaviors

The paper introduces the Behavior-to-Value (B2V) task, which seeks to identify consumer values from e-commerce behavioral trajectories. It presents the E-commerce Consumption Value Taxonomy (ECVT) and the B2V-Bench dataset, derived from anonymized Taobao logs and covering 25 purchase behaviors with associated value orientations. A new model, B2V-Verifier, is proposed to improve value measurement accuracy, achieving a 34% boost in multi-label classification over strong LLM baselines.

By Peixuan Hou, Bin Chen, Li He, Jian Xu, Bo Zheng, Xiuli Ma, Guojie Song
arXiv AI
Sep 25

Beyond Surface Style: Aligning Multi-Turn User Simulators with Behavioral Consistency

The paper introduces TRACER, a multi‑turn user simulator that models evolving user intent and aligns simulated behavior with real interaction trajectories. TRACER is trained first with supervised fine‑tuning on real dialogues and then with reinforcement learning that uses hierarchical outcome‑ and trajectory‑level rewards to address reward sparsity and credit assignment. In real customer‑service sessions, TRACER‑7B outperforms the best baseline by 11.4 conversion F1, achieves the lowest group‑level conversion‑rate error and semantic trajectory distance, and generalizes to out‑of‑distribution scenarios, while human Turing tests show its conversations appear natural. The authors also present the Dynamic Marketing Benchmark, which evaluates both persuasion effectiveness and response quality of large language models through simulated interactions, demonstrating that higher response quality does not always lead to higher conversion rates.

By Geng Chen, Ruotong Pan, Zhirui Yang, Qiqi He, Jiawei Chen, Zhang Yunfei, Chongyuan Chen, Minxuan Lv, Zheng Yang, Win-Bin Huang, Xiangyu Wu, Wenwu Ou
arXiv Computation and Language
Sep 10

See Better, Foresee Better, Act Wiser: Physically Grounded Proactive Modeling and Decision Making

arXiv:2606.03371v4 Announce Type: replace Abstract: Reliable proactive agents must choose an action and judge whether current evidence is sufficient to act. We study retail service from sparse third-...

By Honghui Zhang, Anna Min, Chenmeinian Guo, Yujia Zhang, Yichen Yu, Zezhou Zhang, Guanyu Liu, Yongming Qin, Chongguo Song, Mengyue Yang, Lei Yu, Tianyu Shi