arXiv Computation and Language
The paper introduces a probe‑free method called the Neuron Separability Index (NSI) to assess how individual neurons in Large Language Models (LLMs) distinguish grammatical from ungrammatical constructions. Using linguistic minimal pairs across 68 paradigms and seven checkpoints, the study finds that raw separability peaks earlier for morphological and syntactic distinctions, but after permutation normalization, single‑unit selectivity is sparse, weak, and narrowly tuned, with rare strongly selective "grandmother neurons". Additionally, whole‑vector linear separability, single‑neuron selectivity, and behavioral competence are largely dissociated, and targeted ablations further separate activation selectivity from causal reliance.
The paper introduces a probe‑free method called the Neuron Separability Index (NSI) to assess how individual neurons in large language models distinguish grammatical from ungrammatical sentences using linguistic minimal pairs. Across 68 linguistic paradigms and seven model checkpoints, the study finds that while raw separability for morphology and syntax peaks early, single‑unit selectivity is sparse and weak, with rare strongly selective "grandmother neurons." Moreover, the research shows a dissociation between whole‑vector linear separability, single‑neuron selectivity, and behavioral competence, and demonstrates that targeted ablations can further separate activation selectivity from causal reliance.
The paper investigates how language models avoid overgeneralizations by distinguishing between two types of indirect negative evidence: preemption and entrenchment. Through controlled rearing experiments on models trained on child‑caregiver conversations, the authors find that models do not exhibit verb‑specific preemption but show weak abstract preemption. Analysis of training dynamics suggests that competing structures act as indirect positive evidence rather than negative in the verb‑specific condition.
By Yixuan Wang, Freda Shi, Kanishka Misra
The paper investigates the intrinsic dimension (ID) of large language model (LLM) representations as an indicator of linguistic complexity. By comparing ID across model layers for coordination vs. subordination, right‑branching vs. center‑embedding, and unambiguous vs. ambiguous attachment, the authors find consistent ID differences that align with established complexity contrasts. Experiments across six LLMs, including representational similarity and layer pruning analyses, confirm that more complex phenomena produce higher ID profiles, with peaks occurring at different layers for each contrast.
By Marco Baroni, Emily Cheng, Iria de-Dios-Flores, Francesca Franzon
arXiv:2602. 09992v2 Announce Type: replace-cross Abstract: Several recent contributions have evaluated the Poverty of the Stimulus Hypothesis (PoSH) using Artificial Neural Networks (ANNs).
By Xiulin Yang, Arianna Bisazza, Nathan Schneider, Ethan Gotlieb Wilcox
arXiv:2609.14384v1 Announce Type: new
Abstract: What must a neural system be capable of to implement language? Current research annotates stimuli with linguistic variables and tests which electrodes,...
By Elliot Murphy
arXiv:2608. 08159v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly reported to exhibit human-like neural and cognitive signatures, including concept cells, mental number lines, and cognitive maps.
By Yuqi Wu, Shengming Zhao, Jie Chen
arXiv:2608.29034v1 Announce Type: cross
Abstract: A wide range of methods have been proposed for interpreting language models, delivering important insights into their inner workings. However, differ...
By Zhang Enyan, R. Thomas McCoy
arXiv:2605. 03058v2 Announce Type: replace-cross Abstract: A central goal of explainable AI is to express large language model (LLM) decision logic symbolically and ground it in internal mechanisms.
By Francesco Sovrano, Gabriele Dominici, Marc Langheinrich
arXiv:2605.30381v2 Announce Type: replace-cross
Abstract: When a language model is fine-tuned to produce systematically incorrect responses, does this training leave a structured, linearly recoverabl...
By Vahideh Zolfaghari
arXiv:2607. 25907v1 Announce Type: cross Abstract: Activation steering controls model behavior by editing internal activations at inference time.
By Deepanshu Mody, Samarth Agarwal, Utkarsh Mittal, Dipesh Mahato
arXiv:2608. 10214v1 Announce Type: new Abstract: Do large language models contain domain-specific parametric shells: concentrated, causally necessary neuron populations whose removal selectively degrades a target domain while sparing others?
By Marcus Armstrong, Navid Ayoobi, Arjun Mukherjee
arXiv:2608.27813v1 Announce Type: new
Abstract: Structural probes were introduced by Hewitt and Manning to reconstruct syntactic trees from a neural language model's latent representations. They are...
By Juan Pablo Vigneaux, Mary Kennedy, Khalil Iskarous, Robert Frank, Matilde Marcolli