arXiv Machine Learning By Xiang Guan, Roger D. Newman-Norlund, Yong Yang, Saeed Ahmadi, Regan Willis, Nadra Salman, Kalil Warren, Srihari Nelakuditi, Chris Rorden, Leonardo Bonilha, Julius Fridriksson

Perturbation-based Regional Interpretability through Subtraction Mapping (PRISM): naming-error dissociations in language models and post-stroke aphasia

Read the original on arXiv Machine Learning →

arXiv:2608. 12717v1 Announce Type: new Abstract: Mechanistic interpretability of large language models lacks spatially resolved, falsifiable tools for testing whether internal components are specialized for distinct cognitive operations.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Aug 10

Recovering Lesion Parameters from Aphasic Picture Naming Error Profiles in Large Language Models

arXiv:2608. 06429v1 Announce Type: cross Abstract: Interpretability methods for large language models (LLMs) describe internal state but do not directly test whether that state is causally sufficient to produce the observed behavior.

By Yong Yang, Roger Newman-Norlund, Xiang Guan, Saeed Ahmadi, Regan Willis, Nadra Salman, Kalil Warren, Sophie Arheix-Parras, Srihari Nelakuditi, Leonardo Bonilha, Christopher Rorden, Rutvik H. Desai, Julius Fridriksson
arXiv AI
Jul 14

Lesioned Multimodal Language Models Reproduce Aphasic Picture-Naming Patterns

arXiv:2607. 11621v1 Announce Type: new Abstract: Aphasia following stroke commonly produces systematic naming errors with characteristic profiles, but whether general-purpose language models not designed for clinical simulation can reproduce these patterns remains untested.

By Yong Yang, Xiang Guan, Sophie Arheix-Parras, Saeed Ahmadi, Roger Newman-Norlund, Leonardo Bonilha, Christopher Rorden, Julius Fridriksson, Rutvik H. Desai, Srihari Nelakuditi
arXiv Computation and Language
4d ago

Cognitive Expert Language Models Better Align with the Corresponding Brain Systems

The study investigates whether language models tailored to specific cognitive domains better align with corresponding brain systems. By prompting and fine‑tuning large language models into six domain experts—sensory, spatial, numerical, reasoning, social, and abstract—the authors find that each expert’s representations more closely match the brain region associated with its domain than other experts. This domain‑specific alignment holds across multiple base models and fMRI datasets, while overall prediction accuracy remains largely unchanged, indicating that regional alignment can be obscured when summarizing across the brain.

By Zhivar Sourati, Mengxuan Helen Wu, Nona Ghazizadeh, Jonas Kaplan, Morteza Dehghani, Samuel A. Nastase
arXiv Machine Learning
Sep 23

Spectral Tail Interventions in Decoder-Only Language Models: Reasoning-Sensitive Weight Structure from Controlled Surgery

The paper investigates how the upper spectral tails of weight matrices in decoder‑only transformer language models influence reasoning behavior. By performing controlled interventions on the query–key product and comparing them to factor‑level surgeries, the authors find that edits targeting the spectral tail more strongly affect model performance across multiple checkpoints and reasoning benchmarks. The study also explores how inverse participation ratios predict accuracy transitions and shows that tail‑aware low‑rank adaptations converge faster than standard methods.

By Ibne Farabi Shihab, Sanjida Akhter, Md Najmus Swaqeeb, Abu Sa-Adat Mohamed Moon-Im Al Ahsan, Anuj Sharma