arXiv AI
3d ago

The Unsampled Truth: Quantifying Prompt Artifacts in LM Psychometrics

The study investigates how different prompt components affect language model responses in psychometric tests. By crossing five distinct baseline personas with five variants of each prompt element—persona wording, task instruction, item wording, and option symbol—the authors measure response shifts using the 1‑Wasserstein distance. Their analysis of 13 small open‑weight language models on the Big Five Inventory and Short Dark Triad reveals that task instruction and option symbol changes often cause more variation than paraphrasing the persona or item, with prompt artifacts explaining over 50% of the variation for many items.

By Nils Schwager, Christoph Hau, Simon M\"unker, Achim Rettinger