arXiv AI

Psychological Competence as a Missing Dimension in AI Evaluation

arXiv:2607. 08285v1 Announce Type: new Abstract: Current AI evaluation frameworks focus primarily on technical performance, including accuracy, robustness, reasoning ability, and policy compliance.

arXiv AI
Jul 24

HARP: The Human--AI Research Platform

arXiv:2607. 20773v1 Announce Type: cross Abstract: Large language models (LLMs) have shifted human--computer interaction from `traditional'' interface journeys toward more conversational exchanges.

By Zeshu Zhu, Natalie Friedman, Kevin Weatherwax, Emily Eiben
arXiv AI
Aug 19

Beyond BFI: The CSI for Enhanced Reliability and Validity in Evaluating LLM Personality Traits

The paper introduces the Core Sentiment Inventory (CSI), a new personality trait evaluation tool for large language models (LLMs) that addresses reliability and validity issues found in existing methods like the Big Five Inventory (BFI). CSI is designed specifically for LLMs, supports both English and Chinese, and provides detailed psychological portraits of model behavior. Experiments show that CSI captures nuanced behavioral patterns, improves reliability, and correlates strongly (above 0.85) with real-world LLM outputs.

By Huanhuan Ma, Haisong Gong, Xiaoyuan Yi, Xing Xie, Philip S. Yu, Dongkuan Xu
arXiv AI
Sep 12

Defining AI Agents: A Compendium of Criteria, Metrics, and Benchmarks

The article "Defining AI Agents: A Compendium of Criteria, Metrics, and Benchmarks" surveys the lack of a standard definition for AI agents and organizes this ambiguity into five dimensions: environmental interaction, learning and adaptation, autonomy, goal‑directed behavior, and temporal coherence. It reviews how each dimension has been conceptualized in prior work and compiles the metrics, benchmarks, and evaluation frameworks used to assess them. The authors also introduce the Agent Compendium, a public digital resource that extends these evaluation methods, aiming to provide a common structure for evaluating and comparing agent capabilities across AI systems.

By Mia Lassiter, Brinnae Bent