arXiv AI By Liangwei Yang, Jielin Qiu, Zixiang Chen, Ming Zhu, Juntao Tan, Zhiwei Liu, Wenting Zhao, Zhujun Lan, Akshara Prabhakar, Silvio Savarese, Huan Wang, Shelby Heinecke

BehaviorBench: Modeling Real-World User Decisions from Behavioral Traces

Read the original on arXiv AI →

arXiv:2606. 02798v1 Announce Type: new Abstract: Many decision-support settings require systems that adapt to individual users, but evaluation data for this problem remain limited.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 12

UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs

arXiv:2608. 10042v1 Announce Type: cross Abstract: Tool-use LLMs are increasingly asked to act on users' behalf, but existing benchmarks usually focus on profile recall, style imitation, generic tool use, or response-level personalization.

By Xuexiong Yin, Zechuan Chen, Yongsen Zheng, Yuxiang Zhang, Jingyuan Yang, Bin Wang, Yubin Wang, Keze Wang
arXiv AI
Sep 23

Are LLMs Good Financial User Simulators? Multi-view Investor Logic Alignment (MILA)

The study examines whether large language models (LLMs) can accurately simulate individual financial users by conducting a longitudinal paper‑trading experiment with 80 participants. Using a rolling next‑day prediction protocol, the researchers compared LLM predictions to a simple recent‑activity persistence baseline across multiple behavioral fidelity levels, from trade occurrence to asset selection and portfolio outcomes. Results show that no LLM consistently outperforms the baseline, with fidelity decreasing at finer behavioral granularity, and that recent trading history largely drives activity predictions while asset selection depends more on available evidence.

By Jiajie He, Jiangyuan Hong, Xintong Chen, Dongling Ni, Wenjin Liu
arXiv AI
Jul 9

Large Behavior Model: A Promptable Digital Twin of the Retail Customer

arXiv:2607. 06993v1 Announce Type: new Abstract: Customer behavior modeling underpins recommendation, marketing, and decision support, yet existing approaches either optimize predictive accuracy without explaining decisions or simulate users without grounding them in real behavioral data.

By Wachiravit Modecrua, Krittin Pachtrachai, Touchapon Kraisingkorn
arXiv AI
Jun 18

Improve Large Language Model Systems with User Logs

arXiv:2602. 06470v3 Announce Type: replace-cross Abstract: Scaling training data and model parameters has long driven progress in large language models (LLMs), but this paradigm is increasingly constrained by the scarcity of high-quality data and diminishing returns from rising computational costs.

By Changyue Wang, Weihang Su, Qingyao Ai, Xingzhao Yue, Rui Zhang, Xiaojia Chang, Yiqun Liu