Measurement Validity in LLM Cultural Alignment
arXiv:2608.29266v1 Announce Type: cross Abstract: Researchers increasingly treat LLM survey responses as a proxy for human cultural values. This includes projecting model outputs onto instruments lik...
arXiv:2608. 06085v1 Announce Type: new Abstract: Survey-country metadata can improve an LLM's forecast of an individual response when informative, yet the same cue may redirect the forecast when assigned at random.
arXiv:2608.29266v1 Announce Type: cross Abstract: Researchers increasingly treat LLM survey responses as a proxy for human cultural values. This includes projecting model outputs onto instruments lik...
Using large language models (LLMs) to simulate diverse human populations has the potential to transform many aspects of computational social science, yet many evaluations score the average response ra...
arXiv:2609.15849v1 Announce Type: cross Abstract: Can LLMs reason through new information like humans, or do they merely retrieve cached opinions? This is critical for silicon sampling, where LLM per...
arXiv:2607. 20441v1 Announce Type: cross Abstract: Every information ecosystem produces beliefs that shape strategic decisions.
Large language models are increasingly used as decision aids whose probability judgments shape downstream choices. Whether those judgments carry a systematic directional tilt has been hard to detect: calibration metrics aggregate unsigned errors, and naturalistic uncertainty offers no ground-truth probability.
arXiv:2608. 19760v1 Announce Type: cross Abstract: Audited against causal ground truth from executed replay in a single-agent tool environment (ALFWorld), none of the step-level credit signals used to train LLM agents -- LLM-judge scores, outcome-conditioned logprob ratios, or the policy's own confidence -- identifies which steps causally matter better than chance.
arXiv:2512. 23847v2 Announce Type: replace-cross Abstract: We develop a statistical procedure to detect lookahead bias in economic forecasts generated by large language models (LLMs).
Prior work on human label variation (HLV) in natural language inference (NLI) has often relied on re-annotation resources that select items by disagreement level. An earlier study (arXiv:2607.
arXiv:2608.18768v2 Announce Type: replace Abstract: Large language models are widely used to simulate survey respondents, yet their outputs are homogeneous and unfaithful to real inter-group differen...
arXiv:2607. 23519v1 Announce Type: cross Abstract: Political audits of large language models (LLMs) usually reduce each to one point on a political compass.
The paper investigates whether confidence scores from a black-box decision model, Jev, truly reflect missing knowledge. Using over 15 public datasets and 6 synthetic task families, the authors find that while Jev’s confidence is calibrated on familiar closed-choice tasks, it fails to indicate when the model lacks relevant information—assigning high confidence to salient options even without answer-relevant data and overestimating accuracy on news beyond its knowledge boundary. Targeted yes/no questions about whether an outcome is settled or whether evidence suffices provide sharper indicators of knowledge gaps, but only when surface cues are controlled.
arXiv:2606. 28963v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to simulate social survey responses, yet their outputs exhibit systematic biases: marginal distributions are skewed, response variance is poorly calibrated, and predictor-outcome relationships are attenuated.