arXiv AI By Bianca Raimondi, Maurizio Gabbrielli

Mechanistic Interpretability of Cognitive Complexity in LLMs via Linear Probing using Bloom's Taxonomy

Read the original on arXiv AI →

arXiv:2602. 17229v2 Announce Type: replace Abstract: The black-box nature of Large Language Models necessitates novel evaluation frameworks that transcend surface-level performance metrics.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Aug 31

Tracing the complexity profiles of different linguistic phenomena through the intrinsic dimension of LLM representations

The paper investigates the intrinsic dimension (ID) of large language model (LLM) representations as an indicator of linguistic complexity. By comparing ID across model layers for coordination vs. subordination, right‑branching vs. center‑embedding, and unambiguous vs. ambiguous attachment, the authors find consistent ID differences that align with established complexity contrasts. Experiments across six LLMs, including representational similarity and layer pruning analyses, confirm that more complex phenomena produce higher ID profiles, with peaks occurring at different layers for each contrast.

By Marco Baroni, Emily Cheng, Iria de-Dios-Flores, Francesca Franzon
arXiv AI
Sep 21

Decoupling Internal Representational Changes and Causal Importance in Fine-Tuned Large Language Models

Fine‑tuning reshapes internal representations of large language models, affecting attention patterns and layer‑wise activations. The study shows that components identified by EAP as important for task performance cluster in specific layers, yet these layers do not align with those undergoing the largest representational changes. Additionally, overlapping EAP components across different tasks do not guarantee cross‑task transfer and can even degrade performance when tasks differ in nature.

By Lingfang Li, Procheta Sen, Shubham Das, Danushka Bollegala
arXiv AI
Sep 18

A Neuropsychologically Grounded Evaluation of LLM Cognitive Abilities

The paper introduces NeuroCognition, a benchmark based on three neuropsychological tests—Raven's Progressive Matrices, Spatial Working Memory, and the Wisconsin Card Sorting Test—to evaluate foundational cognitive abilities in large language models (LLMs). It finds that while LLMs excel on text tasks, their performance drops on image-based and more complex tasks, and they fail different parts of the same tasks compared to humans. NeuroCognition correlates with standard general-capability benchmarks yet measures distinct cognitive skills, highlighting where LLMs align with or diverge from human-like intelligence.

By Faiz Ghifari Haznitrama, Faeyza Rishad Ardi, Alice Oh