arXiv:2608. 05246v1 Announce Type: new Abstract: Existing personalized LLM benchmarks primarily rely on textual personas or isolated behavioral signals, providing limited evaluation of cross-domain behavioral personalization, where responses must be grounded in heterogeneous daily-life activities.
By Jiahao Zhang, Yongzhi Tong, Zelin Fu, Pengde Zhao, Yanmei Jiang, Jiang Feng, Min Yang
arXiv:2609.01257v1 Announce Type: new
Abstract: As LLM-based human simulators are increasingly used for policy, evaluation, and training, they must faithfully reproduce real behavioral patterns. Whil...
By Yi Fei Cheng, Fan Yang, Iremsu Bas, Koichiro Niinuma, Narishige Abe, David Lindlbauer
arXiv:2609.07594v1 Announce Type: cross
Abstract: Conversational Task Assistants (CTAs) are multimodal dialogue systems that support users in complex real-world tasks such as cooking and DIY through...
By Rafael Ferreira, Diogo Tavares, Diogo Gl\'oria-Silva, David Semedo, Jo\~ao Magalh\~aes
arXiv:2607. 27056v1 Announce Type: new Abstract: Personalized agents are increasingly applied to assist users across a wide range of tasks.
By Lingyang Zeng, Guangze Chen, Kaichen Yu, Zhicheng Pan, Siyang Weng, Zirui Hu, Xiangyun Du, Hailin He, Rong Zhang, Chengcheng Yang, Kai Huang, Xuan Zhou
PersonaMem-v3 is a benchmark and evaluation harness designed to assess omni-platform personal intelligence for AI agents. It is built from over one million anonymized real-world engagement histories, covering social media, chatbots, calendars, and AI companions, and tracks user preferences and habits over time. The benchmark tests agents on personalization, LLM-powered recommendation, proactiveness, agentic tool use, and geo-temporal reasoning, evaluating their ability to infer holistic user understanding, personalize responses, rerank recommendations, follow user steering, and avoid inappropriate personalization.
By Bowen Jiang, Yuan Yuan, Zhuoqun Hao, Yuchen Liu, Maohao Shen, Sihao Chen, Gregory Wornell, Chris Callison-Burch, Lyle Ungar, Dan Roth, Qi Guo, Xiangjun Fan, Camillo J. Taylor, Hanchao Yu
arXiv:2607. 20734v1 Announce Type: new Abstract: As LLMs become more capable, they are increasingly deployed as collaborative agents, taking on user-delegated tasks through iterative interaction.
By Jihoon Tack, Philippe Laban, Jennifer Neville
The paper investigates whether human psychometric questionnaires can reliably characterize large language models (LLMs) in everyday interactions. By comparing eight open‑source LLMs’ value and personality profiles from Likert self‑reports (PVQ‑40/21 and BFI‑44/10) with generation probabilities on value‑laden user queries, the authors find substantial divergence between the two methods. The study shows that questionnaire items contain explicit lexical cues that lead models to respond in socially desirable ways, whereas realistic user queries lack such cues, and demographic persona prompts shift questionnaire responses but not generation outputs, indicating that questionnaire scores overestimate LLMs’ true behavioral tendencies.
By Woojung Song, Dongmin Choi, Yoonah Park, Jongwook Han, Eun-Ju Lee, Yohan Jo
arXiv:2606. 16023v1 Announce Type: new Abstract: Human mobility appears highly diverse, yet much of a person's daily mobility can be explained by a small set of recurring behavioral templates, such as commuting, school-centered activities, caregiving, nightlife, or errand patterns.
By Bita Azarijoo, John Krumm, Cyrus Shahabi
arXiv:2609.18273v1 Announce Type: new
Abstract: Web browsing often appears ephemeral: users visit a few websites, complete a task, and move on. However, even short fragments of browsing activity can...
By Ralph Elsaghbini, Omran Berjawi, Walid Fahs, Rida Khatoun
arXiv:2608. 02491v2 Announce Type: replace Abstract: Language models have taken on the role of a very new type of technology, by virtue of their "human-ness" and rapid integration into users' daily lives.
By Nicole Mitchell, Dhruv Agarwal, Maty Bohacek, Remi Denton, Roma Patel
arXiv:2609.14849v1 Announce Type: cross
Abstract: We characterize how people are turning to LLMs as oracles: all-knowing authorities on subjective personal questions. Motivated by risks to users' aut...
By Myra Cheng, Lujain Ibrahim, Grace Liu, Michelle S. Lam, Vishakh Padmakumar, Nick Madibekov, Diyi Yang, Dan Jurafsky
arXiv:2608. 12582v1 Announce Type: cross Abstract: AI journaling tools can tailor prompts to a person's own sensed behavior, but it is unclear which behaviors respond to them.
By Nadia Mehjabin, Henry Kautz, Subigya Nepal