arXiv AI

Apparent Psychological Profiles of Large Language Models are Largely a Measurement Artifact

arXiv:2606. 20205v1 Announce Type: new Abstract: Psychological instruments designed for humans are increasingly used to assign large language models (LLMs) stable psychological profiles that affect their usability, safety assessment, and use as proxies for human participants in research.

arXiv AI
Jul 10

Persona Cartography: Charting Language Model Personality Traits in Weight Space

arXiv:2607. 07916v1 Announce Type: new Abstract: Large language models exhibit recurring behavioural patterns -- personas -- that shape generalisation and safety, but we lack reliable tools for decomposing, measuring, and controlling them.

By Luke Baines, Anton Gonzalvez Hawthorne, Mariia Koroliuk, Irakli Shalibashvili, Cl\'ement Dumas, Konstantinos Voudouris, David Demitri Africa
arXiv Computation and Language
Sep 22

Measuring Behavioural Signatures of Large Language Models through Psychometric Profiling

arXiv:2609.22934v1 Announce Type: new Abstract: Large language models (LLMs) increasingly mediate human decisions and communication, yet their behavioural regularities remain difficult to characteriz...

By Yu Sha, Junqi Tao, Dixin Zhou, Yansheng Tu, Mingyang Chen, Xiang Fan, Yang Liu, Mengquan Yang, Jie Lin, Jiahui Fu, Hua Zheng, Benwei Zhang, Zhou Kai
arXiv AI
Sep 4

Human Psychometric Questionnaires Mischaracterize LLM Behavior

The paper investigates whether human psychometric questionnaires can reliably characterize large language models (LLMs) in everyday interactions. By comparing eight open‑source LLMs’ value and personality profiles from Likert self‑reports (PVQ‑40/21 and BFI‑44/10) with generation probabilities on value‑laden user queries, the authors find substantial divergence between the two methods. The study shows that questionnaire items contain explicit lexical cues that lead models to respond in socially desirable ways, whereas realistic user queries lack such cues, and demographic persona prompts shift questionnaire responses but not generation outputs, indicating that questionnaire scores overestimate LLMs’ true behavioral tendencies.

By Woojung Song, Dongmin Choi, Yoonah Park, Jongwook Han, Eun-Ju Lee, Yohan Jo
arXiv AI
Sep 1

The Unsampled Truth: Quantifying Prompt Artifacts in LM Psychometrics

The study investigates how different prompt components affect language model responses in psychometric tests. By crossing five distinct baseline personas with five variants of each prompt element—persona wording, task instruction, item wording, and option symbol—the authors measure response shifts using the 1‑Wasserstein distance. Their analysis of 13 small open‑weight language models on the Big Five Inventory and Short Dark Triad reveals that task instruction and option symbol changes often cause more variation than paraphrasing the persona or item, with prompt artifacts explaining over 50% of the variation for many items.

By Nils Schwager, Christoph Hau, Simon M\"unker, Achim Rettinger