arXiv AI

Human Psychometric Questionnaires Mischaracterize LLM Behavior

The paper investigates whether human psychometric questionnaires can reliably characterize large language models (LLMs) in everyday interactions. By comparing eight open‑source LLMs’ value and personality profiles from Likert self‑reports (PVQ‑40/21 and BFI‑44/10) with generation probabilities on value‑laden user queries, the authors find substantial divergence between the two methods. The study shows that questionnaire items contain explicit lexical cues that lead models to respond in socially desirable ways, whereas realistic user queries lack such cues, and demographic persona prompts shift questionnaire responses but not generation outputs, indicating that questionnaire scores overestimate LLMs’ true behavioral tendencies.

arXiv Computation and Language
5d ago

PACIFIC: Can LLMs Discern the Psychometric Traits Influencing Your Preferences? Personality-Driven Preference Alignment in LLMs

PACIFIC is a framework that aligns large language model responses with user preferences by leveraging stable Big‑Five personality traits as a latent signal. The authors built a 1,200‑pair dataset covering diverse domains and trait directions, and found that trait‑aligned contexts enable LLMs to achieve near‑ceiling accuracy (up to 99%) in personalized QA. They also introduced a persona‑aware contrastive retriever (PiRAG) that improves label‑free accuracy from 30% to 43% over standard semantic retrieval, highlighting retrieval as the main bottleneck.

By Tianyu Zhao, Siqi Li, Yasser Shoukry, Salma Elmalaki
arXiv AI
Aug 19

Beyond BFI: The CSI for Enhanced Reliability and Validity in Evaluating LLM Personality Traits

The paper introduces the Core Sentiment Inventory (CSI), a new personality trait evaluation tool for large language models (LLMs) that addresses reliability and validity issues found in existing methods like the Big Five Inventory (BFI). CSI is designed specifically for LLMs, supports both English and Chinese, and provides detailed psychological portraits of model behavior. Experiments show that CSI captures nuanced behavioral patterns, improves reliability, and correlates strongly (above 0.85) with real-world LLM outputs.

By Huanhuan Ma, Haisong Gong, Xiaoyuan Yi, Xing Xie, Philip S. Yu, Dongkuan Xu
arXiv Computation and Language
Sep 3

When Persona Attributes Improve Population Alignment in Large Language Models

The paper investigates how persona prompting—using short textual descriptions of individuals—to align large language models (LLMs) with human survey responses. It examines the impact of selecting different persona attributes and finds that not all attribute combinations improve performance, suggesting that the variation in human responses to survey questions may explain mixed results. The study evaluates multiple attribute selection methods across four social surveys, two countries, six LLMs, and twenty prediction tasks, offering guidance on when persona prompting is beneficial and which attribute choices are most effective.

By Leon Fr\"ohling, Jens Rupprecht, Markus Strohmaier, Claudia Wagner