arXiv AI By Zhicheng Lin

From Prompts to Constructs: A Dual-Validity Framework for Large Language Model Research in Psychology

Read the original on arXiv AI →

arXiv:2506. 16697v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are entering psychological research both as tools and as objects of inquiry.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 19

Beyond BFI: The CSI for Enhanced Reliability and Validity in Evaluating LLM Personality Traits

The paper introduces the Core Sentiment Inventory (CSI), a new personality trait evaluation tool for large language models (LLMs) that addresses reliability and validity issues found in existing methods like the Big Five Inventory (BFI). CSI is designed specifically for LLMs, supports both English and Chinese, and provides detailed psychological portraits of model behavior. Experiments show that CSI captures nuanced behavioral patterns, improves reliability, and correlates strongly (above 0.85) with real-world LLM outputs.

By Huanhuan Ma, Haisong Gong, Xiaoyuan Yi, Xing Xie, Philip S. Yu, Dongkuan Xu
arXiv AI
Aug 6

Item Response Theory for AI Safety

arXiv:2608. 05086v1 Announce Type: new Abstract: Language models differ in how safely they behave and these differences are measured by safety benchmarks.

By Joshua Fonseca Rivera (Independent), Neil Shah (Independent), David Demitri Africa (UK AI Security Institute), Konstantinos Voudouris (UK AI Security Institute)
arXiv AI
Sep 17

Cultural Competence in Context: A Large Language Model Passes the Turing Test in Finland

The study reports that ChatGPT 5.2 successfully passed a Finnish-language Turing Test conducted in Finland, contrary to expectations that uneven representation of Finnish in training data would hinder performance. The authors attribute failures mainly to participants using colloquial Finnish cues to identify human authorship. They propose reinterpreting the Turing Test as a comparative method to assess an AI’s credible membership in a specific social world, highlighting its utility for probing the human‑machine boundary across domains.

By Otto Segersven, Pentti Henttonen
Hugging Face Trending Papers
Jul 13

When the Target Domain Changes: AI-Mediated Construct Drift in High-Stakes English Language AssessmenW

High-stakes English proficiency tests treat standardized, unaided performance as evidence for score interpretations about academic English proficiency. This interpretation remains meaningful, but as target language use domains increasingly involve generative AI, the extrapolation from unaided test performance to academic communicative readiness becomes less self-evident.