arXiv:2507. 04491v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are rapidly being integrated into psychological and behavioral research as research tools, evaluation targets, human simulators, and cognitive models.
By Zhicheng Lin
arXiv:2607. 20773v1 Announce Type: cross Abstract: Large language models (LLMs) have shifted human--computer interaction from `traditional'' interface journeys toward more conversational exchanges.
By Zeshu Zhu, Natalie Friedman, Kevin Weatherwax, Emily Eiben
The paper introduces the Core Sentiment Inventory (CSI), a new personality trait evaluation tool for large language models (LLMs) that addresses reliability and validity issues found in existing methods like the Big Five Inventory (BFI). CSI is designed specifically for LLMs, supports both English and Chinese, and provides detailed psychological portraits of model behavior. Experiments show that CSI captures nuanced behavioral patterns, improves reliability, and correlates strongly (above 0.85) with real-world LLM outputs.
By Huanhuan Ma, Haisong Gong, Xiaoyuan Yi, Xing Xie, Philip S. Yu, Dongkuan Xu
arXiv:2510. 12201v2 Announce Type: replace Abstract: As AI becomes more common in everyday living, there is an increasing demand for intelligent systems that are both performant and understandable.
By Aline Mangold, Juliane Zietz, Susanne Weinhold, Sebastian Pannasch
arXiv:2509.08494v2 Announce Type: replace-cross
Abstract: As humans delegate more tasks and decisions to artificial intelligence (AI), we risk losing control of our individual and collective futures....
By Benjamin Sturgeon, Daniel Samuelson, Jacob Haimes, Jacy Reese Anthis
arXiv:2608. 05710v1 Announce Type: new Abstract: When an AI system is deployed, the individuals who use and or are evaluated by it form beliefs about how the system operates and use those beliefs to strategically present their preferences, behaviors, or attributes.
By Keziah Naggita
arXiv:2607. 01034v1 Announce Type: cross Abstract: Large language model (LLM)-based conversational agents (CAs) are now ubiquitous, creating new opportunities for AI-mediated behavior change.
By Hasibur Rahman, Smit Desai
arXiv:2605. 28591v2 Announce Type: replace-cross Abstract: The validity of AI safety evaluations depends on models behaving consistently across controlled and deployment settings.
By Katharina Deckenbach, Haritz Puerto, Jonas Geiping, Sahar Abdelnabi
arXiv:2506. 16697v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are entering psychological research both as tools and as objects of inquiry.
By Zhicheng Lin
arXiv:2607. 25057v1 Announce Type: new Abstract: As conversational AI systems become increasingly integrated into daily life, their potential effects on user well-being require ongoing attention.
By Jina Suh, Mihaela Vorvoreanu, Forough Poursabzi-Sangdeh, Emily Tseng, Eugenia Kim, Luke Nicholls, James W. Pennebaker, Eric Horvitz
The article "Defining AI Agents: A Compendium of Criteria, Metrics, and Benchmarks" surveys the lack of a standard definition for AI agents and organizes this ambiguity into five dimensions: environmental interaction, learning and adaptation, autonomy, goal‑directed behavior, and temporal coherence. It reviews how each dimension has been conceptualized in prior work and compiles the metrics, benchmarks, and evaluation frameworks used to assess them. The authors also introduce the Agent Compendium, a public digital resource that extends these evaluation methods, aiming to provide a common structure for evaluating and comparing agent capabilities across AI systems.
By Mia Lassiter, Brinnae Bent
LLM companions are deployed at scale in personally consequential settings, yet poorly evaluated. Existing benchmarks use hand-authored scenarios and prompted simulators, aggregate empathy into one score, and overlook judge biases such as same-family favoritism and scale drift.