arXiv AI

Psychological Competence as a Missing Dimension in AI Evaluation

arXiv:2607. 08285v1 Announce Type: new Abstract: Current AI evaluation frameworks focus primarily on technical performance, including accuracy, robustness, reasoning ability, and policy compliance.

arXiv AI
Jul 24

HARP: The Human--AI Research Platform

arXiv:2607. 20773v1 Announce Type: cross Abstract: Large language models (LLMs) have shifted human--computer interaction from `traditional'' interface journeys toward more conversational exchanges.

By Zeshu Zhu, Natalie Friedman, Kevin Weatherwax, Emily Eiben