arXiv AI

Do Assessment Instruments Measure the Same Thing for Humans and LLMs? A Latent Structure Analysis

arXiv:2608. 15630v1 Announce Type: cross Abstract: The rapid development and growing deployment of large language models (LLMs) have made it increasingly important to understand their capabilities.

arXiv AI
Aug 19

Interpretable Humans, Alien LLMs: Expert Analysis of Latent Structures in Assessment Responses

The study examines whether the latent factors that explain large language model (LLM) performance correspond to human‑interpretable cognitive constructs. Using exploratory factor analysis on responses from humans and six LLMs in quantitative reasoning and chemistry, subject‑matter experts could interpret most human‑derived factors but struggled to ascribe meaning to LLM‑derived factors, especially in quantitative reasoning and only partially in chemistry. The results suggest that LLMs often rely on statistically opaque mechanisms distinct from human reasoning.

By Alona Strugatski, Licol Zeinfeld, Jason Cooper, Shelley Rap, Gil Schwarts, Giora Alexandron