arXiv:2606. 24162v1 Announce Type: cross Abstract: Foundation models have been increasingly applied to behavioral science domains such as psychology, sociology, and economics.
By Jin Huang, Yutong Xie, Wanli Song, Xingjian Zhang, Walter Yuan, Matthew O. Jackson, Qiaozhu Mei
PersonaForge is a user‑simulation framework that generates realistic multi‑turn interactions between users and agentic systems, addressing the gap that most training data assumes single‑turn queries. It uses a four‑dimensional persona space, SOUL‑driven behavioral control calibrated to real‑user statistics, and Reverse Deep Construction from authentic seed queries to create a 6.3K‑record training set and a 138‑task benchmark called PersonaForge‑Bench across 20 professional domains. Experiments with Qwen3.5‑27B show that training with PersonaForge improves composite scores by 4.1%, especially in Task Completion (+6.0%) and Response Quality (+6.8%), while also reducing turns and tool calls, indicating more efficient interactions.
By Hanglong Lv, Dawei Zhu, Lei Li, Bowen Ye, Huaqiu Liu, Yifan Song, Bofei Gao, Weimin Xiong, Jinhao Dong, Chenhong He, Lingpeng Kong, Qi Liu, Tong Yang, Fuli Luo
arXiv:2608. 12253v1 Announce Type: cross Abstract: Multi-agent reinforcement learning for human-AI interaction typically relies on a single large language model to simulate user behavior.
By Simon Yu, Nicholas Tomlin, Marwa Abdulhai, Ximing Lu, Derek Chong, Abe Hou, Dilara Soylu, Sergey Levine, Christopher D. Manning, Weiyan Shi
arXiv:2609.00608v1 Announce Type: new
Abstract: LLM-based user simulators are increasingly used to evaluate autonomous agents at scale, in place of costly human evaluations. Despite this promise, the...
By Daeheon Jeong, Yoonjoo Lee, Eugene Choi, Sinie van der Ben, Juho Kim
arXiv:2609.15972v1 Announce Type: cross
Abstract: As language models become more capable, long-term collaboration in learning, reasoning, and decision-making calls for a deeper understanding of the p...
By Zixuan Wang, Yufan Zhou, Jinzhou Tang, Xinle Yu, Chengjun Wu, Lyumanshan Ye, Zhaoxiang Feng, Letian Peng, Adyasha Patra, Fan Bai, Enze Ma, Zhengding Hu, Jianyang Gu, Zhao Wang, Yufei Ding, Jingbo Shang, Tianmin Shu, Zhiting Hu, Zhen Wang
arXiv:2604. 19139v3 Announce Type: replace-cross Abstract: As Large Language Models (LLMs) continue to evolve through alignment techniques such as Reinforcement Learning from Human Feedback (RLHF) and Constitutional AI, a growing and increasingly conspicuous phenomenon has emerged: the proliferation of verbal tics--repetitive, formulaic linguistic patterns that pervade model outputs.
By Shuai Wu, Xue Li, Yanna Feng, Yufang Li, Zhijun Wang, Ran Wang
ZenGen is a comprehensive framework designed to develop socially intelligent large language models (LLMs). It introduces SoMBench, a psychology‑grounded benchmark with 3 primary and 17 secondary dimensions, and shows that current LLMs perform poorly on social cognition tasks. ZenGen combines supervised fine‑tuning, on‑policy distillation, and rubric‑based reinforcement learning to improve internalization, while Actio provides deployment‑time grounding with typed supports such as PRISM, Starling, SAGE, and gated RAG, yielding significant performance gains across benchmarks.
By ZenGen Team, Ao Xiang, Bi Jingping, Chen Jiahui, Chen Lehan, Chen Yilin, Cheng Xueqi, Fan Yixing, Gan Kairong, Gao Haowen, Gao Jinhua, Gao Shuxuan, Gong Chang, Guo Jiafeng, Guo Ruijie, Han Zhouyu, He Guangfu, He Yichun, Jiang Shuo, Jing Shaoling, Jing Ya, Lei Chenhao, Lei Yan, Li Anqi, Li Chengao, Li Haoyu, Li Shitian, Liang Xinjian, Liu Zhaoge, Lyu Xingyu, Nie Zhuwei, Pang Liang, Quan Zeping, Shan Shiguang, Shen Huawei, Tang Xinran, Tian Feng, Wang Qian, Wang Ruiping, Wang Xiaohong, Xia Zaiyu, Xiao Yi, Xu Jiayuan, Xu Kehan, Xu Qianqian, Xu Tianyu, Xu Yongjun, Yang Haoming, Yang Jun, Yao Di, Yu Xiaoming, Zhang Futong, Zhang Jie, Zhang Shixuan, Zhang Yuxuan, Zhao Xinyu, Zhao Zhuoran, Zhong Yunfei, Zhu Shengyu
arXiv:2411. 10109v3 Announce Type: replace Abstract: Machine learning can predict human behavior well when substantial structured data are available for well-defined outcomes.
By Joon Sung Park, Carolyn Q. Zou, Jonne Kamphorst, Niles Egan, Aaron Shaw, Benjamin Mako Hill, Carrie Cai, Meredith Ringel Morris, Percy Liang, Robb Willer, Michael S. Bernstein
GAUGE is a new offline protocol that evaluates whether the common practice of using an LLM-as-a-judge to rank task‑oriented agents actually aligns with a verifiable reward. Across 25 agents from six providers on two benchmarks, GAUGE finds that user satisfaction scores are largely uncorrelated with task success, and that the judge’s ranking loses precision when agents are closely matched in performance. The study highlights a gap between ranking validity and construct validity in current evaluation practices.
By Umesh Bodhwani, Thanh Tran, Kai Wei
arXiv:2606. 29495v1 Announce Type: new Abstract: Social influence dialogue changes user behavior by altering internal cognitive states.
By Minghui Ma, Bin Guo, Han Wang, Mengqi Chen, Jingqi Liu, Yan Liu, Zhiwen Yu
arXiv:2603. 03824v2 Announce Type: replace Abstract: Humans often become more self-aware under threat, yet can lose self-awareness when absorbed in a task; we hypothesize that language models exhibit environment-dependent \textit{evaluation awareness}.
By Maheep Chaudhary
KnowSim introduces an evaluation framework that uses a user simulator with explicit knowledge states to assess how well large language models calibrate information to users. The simulator represents knowledge as a graph of Information Units with prerequisite relationships and updates these states based on learning theory. KnowSim computes Knowledge Gain, Delivery Calibration, and Cognitive Overload metrics, and its rankings align with human judgments, outperforming baseline simulators and revealing model performance differences across user knowledge levels.
By Yoonjoo Lee, Hyoungwook Jin, Tae Soo Kim, Shaoyang Zhang, Philippe Laban, Q. Vera Liao