arXiv AI By Alona Strugatski, Licol Zeinfeld, Jason Cooper, Shelley Rap, Gil Schwarts, Giora Alexandron

Interpretable Humans, Alien LLMs: Expert Analysis of Latent Structures in Assessment Responses

Read the original on arXiv AI →

The study examines whether the latent factors that explain large language model (LLM) performance correspond to human‑interpretable cognitive constructs. Using exploratory factor analysis on responses from humans and six LLMs in quantitative reasoning and chemistry, subject‑matter experts could interpret most human‑derived factors but struggled to ascribe meaning to LLM‑derived factors, especially in quantitative reasoning and only partially in chemistry. The results suggest that LLMs often rely on statistically opaque mechanisms distinct from human reasoning.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 16

Do LLMs Have Values? A Quantitative Analysis and Alignment Framework for Values in Large Language Models

The paper investigates whether large language models (LLMs) possess intrinsic value systems and how to quantify and align them. By projecting responses from 106 LLMs and 95,000 human survey profiles into a shared sociological space, the authors confirm that LLMs do have values, though these values form a concentrated, idealized core rather than mirroring human diversity. They introduce the Prior-Environment-Cognition (PEC) framework to mathematically define value expression and propose an adaptive Alignment Prescription that identifies minimal interventions—ranging from prompts to targeted parameter updates—to steer LLM values efficiently without harming general performance.

By Keqing Zhang, Jingyu Chen, Yufan Liu, Yongqiang Zhu, Nai Ding, Lai Jiang, Congyan Lang, Bing Li, Weiming Hu
arXiv Machine Learning
Sep 10

Do Reasoning Representations Help Humans Evaluate LLM Outputs?

The paper investigates whether reasoning representations—explanations for large language model outputs—aid humans in evaluating those outputs. A controlled human study tested six reasoning formats across tasks of varying complexity, measuring structural understanding, error detection, and trust calibration. Results revealed a mismatch: participants favored planning- and decomposition-based representations, yet simpler chain-of-thought traces better supported verification, trust, and interpretability, while preferred formats increased calibration risks.

By Jaewoo Lim, Sungbok Shin, Sanghyun Hong