arXiv:2608. 10042v1 Announce Type: cross Abstract: Tool-use LLMs are increasingly asked to act on users' behalf, but existing benchmarks usually focus on profile recall, style imitation, generic tool use, or response-level personalization.
By Xuexiong Yin, Zechuan Chen, Yongsen Zheng, Yuxiang Zhang, Jingyuan Yang, Bin Wang, Yubin Wang, Keze Wang
arXiv:2606. 24162v1 Announce Type: cross Abstract: Foundation models have been increasingly applied to behavioral science domains such as psychology, sociology, and economics.
By Jin Huang, Yutong Xie, Wanli Song, Xingjian Zhang, Walter Yuan, Matthew O. Jackson, Qiaozhu Mei
The study examines whether large language models (LLMs) can accurately simulate individual financial users by conducting a longitudinal paper‑trading experiment with 80 participants. Using a rolling next‑day prediction protocol, the researchers compared LLM predictions to a simple recent‑activity persistence baseline across multiple behavioral fidelity levels, from trade occurrence to asset selection and portfolio outcomes. Results show that no LLM consistently outperforms the baseline, with fidelity decreasing at finer behavioral granularity, and that recent trading history largely drives activity predictions while asset selection depends more on available evidence.
By Jiajie He, Jiangyuan Hong, Xintong Chen, Dongling Ni, Wenjin Liu
arXiv:2609.38397v1 Announce Type: new
Abstract: Virtual clients offer a cost-effective approach to support applications such as A/B testing, recommender system development, and interface evaluation....
By Yunan Lu, Shuang Xie, Meghna Allamudi, Mingyu Zhao, Han Li, Lingyun Wang, Zhou Yu
arXiv:2607. 06993v1 Announce Type: new Abstract: Customer behavior modeling underpins recommendation, marketing, and decision support, yet existing approaches either optimize predictive accuracy without explaining decisions or simulate users without grounding them in real behavioral data.
By Wachiravit Modecrua, Krittin Pachtrachai, Touchapon Kraisingkorn
arXiv:2602. 06470v3 Announce Type: replace-cross Abstract: Scaling training data and model parameters has long driven progress in large language models (LLMs), but this paradigm is increasingly constrained by the scarcity of high-quality data and diminishing returns from rising computational costs.
By Changyue Wang, Weihang Su, Qingyao Ai, Xingzhao Yue, Rui Zhang, Xiaojia Chang, Yiqun Liu
arXiv:2609.38043v1 Announce Type: new
Abstract: Interactive agent benchmarks and multi-turn reinforcement learning increasingly place a second language model in the role of the user. This simulated u...
By Ashish Jain, Armaan Sandhu
arXiv:2604. 07343v2 Announce Type: replace-cross Abstract: Pluralistic alignment has emerged as a critical frontier in the development of Large Language Models (LLMs), with reward models (RMs) serving as a central mechanism for capturing diverse human values.
By Qiyao Ma, Dechen Gao, Rui Cai, Boqi Zhao, Hanchu Zhou, Junshan Zhang, Zhe Zhao
arXiv:2608. 04095v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly used as personalized assistants in high-stakes domains such as financial advising, yet it remains unclear whether they can maintain and update an individualized user model over long horizons.
By Ben Wang, Kang Zhou, Lifan Guo, Feng Chen, Chi Zhang
Large language models (LLMs) are being tested as simulators of individual financial decision-making. In a controlled paper‑trading experiment with 120 volunteers, the study evaluated whether an LLM could predict a participant’s next‑day trading action, chosen security, and transaction size using only pre‑cutoff information. Results showed that including market context improved predictions of actions and tickers, but sizing remained challenging, and the models tended to over‑predict hold actions, under‑predict sells, and simplify multi‑security trades.
By Jiajie He, Jiangyuan Hong, Dongling Ni, Wenjin Liu, Xintong Chen
HyperTrace is a training‑free framework that personalizes large language models by tracing latent user preferences online. It maintains interpretable natural‑language hypotheses about short‑term intent and long‑term preferences, updating them with an SMC‑style reweighting process driven by an LLM‑based surrogate choice model. Experiments on PRISM and PersonaMem‑v2 demonstrate that HyperTrace improves response alignment, preference prediction, and profile consistency compared to strong online baselines.
By Jianzhi Shen, Keyu Mao, Minghao Shao, Chuanyang Jin, Yusong Wang, Ailiang Lin, Kotaro Funakoshi, Manabu Okumura, Tianmin Shu, Muhammad Shafique
arXiv:2607. 20471v1 Announce Type: new Abstract: Personalization, the act of varying a message to induce action from a specific receiver while keeping sender, channel, and time fixed, has a long tradition in psychology and marketing as a two-party problem in which sender and receiver have independent objectives.
By Ashutosh Srivastava, Siddharth Yedlapati, Vinay Aggarwal, Yaman Kumar Singla, Shashwat Dixit, Jitendra Ajmera, Balaji Krishnamurthy