The paper introduces VIBE‑Bench, a new benchmark designed to test personalized large language models (PLLMs) in a regime where user profile cues and query‑specific preferences do not share the same conceptual space, a situation termed profile‑preference conceptual misalignment (PRCM). VIBE‑Bench contains 3,504 personas, 12,239 dialogues, and a manually verified gold test set, and includes two psychology‑grounded tasks that require cross‑concept preference reasoning beyond surface semantic overlap. Experiments show that existing PLLMs largely depend on shallow semantic correlations and struggle to learn robust cross‑concept mappings, highlighting PRCM as a distinct failure mode for personalization models.
By Yiwen Jiang, Yang Deng, Stephanie Fong, Zimu Wang, Yaling Shen, Wei Feng, Hongxi Yang, Xiangyu Zhao, Zhongxing Xu, Deval Mehta, Xuelian Cheng, Zongyuan Ge
arXiv:2607. 27056v1 Announce Type: new Abstract: Personalized agents are increasingly applied to assist users across a wide range of tasks.
By Lingyang Zeng, Guangze Chen, Kaichen Yu, Zhicheng Pan, Siyang Weng, Zirui Hu, Xiangyun Du, Hailin He, Rong Zhang, Chengcheng Yang, Kai Huang, Xuan Zhou
The paper investigates whether human psychometric questionnaires can reliably characterize large language models (LLMs) in everyday interactions. By comparing eight open‑source LLMs’ value and personality profiles from Likert self‑reports (PVQ‑40/21 and BFI‑44/10) with generation probabilities on value‑laden user queries, the authors find substantial divergence between the two methods. The study shows that questionnaire items contain explicit lexical cues that lead models to respond in socially desirable ways, whereas realistic user queries lack such cues, and demographic persona prompts shift questionnaire responses but not generation outputs, indicating that questionnaire scores overestimate LLMs’ true behavioral tendencies.
By Woojung Song, Dongmin Choi, Yoonah Park, Jongwook Han, Eun-Ju Lee, Yohan Jo
HyperTrace is a training‑free framework that personalizes large language models by tracing latent user preferences online. It maintains interpretable natural‑language hypotheses about short‑term intent and long‑term preferences, updating them with an SMC‑style reweighting process driven by an LLM‑based surrogate choice model. Experiments on PRISM and PersonaMem‑v2 demonstrate that HyperTrace improves response alignment, preference prediction, and profile consistency compared to strong online baselines.
By Jianzhi Shen, Keyu Mao, Minghao Shao, Chuanyang Jin, Yusong Wang, Ailiang Lin, Kotaro Funakoshi, Manabu Okumura, Tianmin Shu, Muhammad Shafique
arXiv:2603.04191v2 Announce Type: replace
Abstract: Large Language Models (LLMs) are increasingly serving as personal assistants, where users may share individual preferences over extended interactio...
By Qianyun Guo, Yibo Li, Yue Liu, Bryan Hooi
arXiv:2608. 11354v1 Announce Type: new Abstract: Modern recommender systems treat observed actions as reliable proxies for user preferences, yet interactions often reflect exploration or comparison rather than stable preference expression.
By Mengyu Chen, Feiyu Lu, Chun-Fu Chen, Lucas Vinh Tran, Jay Katukuri
The paper introduces ReaLMem, a benchmark built from authentic multi‑year personal visual archives with first‑person annotations, designed to evaluate AI systems on factual recall, persona inference, and predictive personalization. It also proposes ChronoProfiler, a temporal‑weighting module that calculates stability scores for user attributes to resolve preference conflicts and enhance personalized decision making. Experiments with multimodal large language models and memory systems show that predictive personalization remains the hardest task, highlight performance gaps, and demonstrate that temporally informed representations significantly improve personalization.
By Wenqi Zhou, Zhuorui Yu, Kaiao Wen, Hao Zheng, Xinyi Zheng, Peiran Wu, Enmin Zhou, Chi-Hao Wu, Junxiao Shen
arXiv:2609.00014v1 Announce Type: cross
Abstract: Persona-driven techniques increasingly adapt large language models (LLMs) to diverse contexts. However, existing methods predominantly rely on rigid,...
By Yuxuan Li, Victor Zhong, Ehsan Kamalloo
The paper introduces the Core Sentiment Inventory (CSI), a new personality trait evaluation tool for large language models (LLMs) that addresses reliability and validity issues found in existing methods like the Big Five Inventory (BFI). CSI is designed specifically for LLMs, supports both English and Chinese, and provides detailed psychological portraits of model behavior. Experiments show that CSI captures nuanced behavioral patterns, improves reliability, and correlates strongly (above 0.85) with real-world LLM outputs.
By Huanhuan Ma, Haisong Gong, Xiaoyuan Yi, Xing Xie, Philip S. Yu, Dongkuan Xu
arXiv:2608. 11493v1 Announce Type: new Abstract: Traditional offline recommendation evaluation relies heavily on complex, manually maintained feature pipelines that are difficult to scale.
By Alireza S. Ziabari, Kat Ellis, Colleen Chan, Ding Tong
arXiv:2606. 06614v1 Announce Type: cross Abstract: Despite growing interest, most evaluations of large language models' (LLMs') personalization abilities have relied on synthetic data.
By Lechen Zhang, Jiarui Liu, Tal August
arXiv:2510. 22170v3 Announce Type: replace Abstract: Persona conditioning is widely used to steer large language model (LLM) behavior, but it is unclear whether it induces stable behavioral structure or superficial variation.
By Alexandra Yost, Shreyans Jain, Shivam Raval, Grant Corser, Allen Roush, Nina Xu, Jacqueline Hammack, Ravid Shwartz-Ziv, Amirali Abdullah