arXiv AI

Localizing Anchoring Pathways in Language Models

arXiv:2606. 12818v1 Announce Type: cross Abstract: Irrelevant numbers in a prompt can shift language model judgments, producing anchoring effects in numerical reasoning.

arXiv Machine Learning
Jun 2

Correcting Gradient-Based Circuit Localization via Interaction-Aware Backpropagation

arXiv:2505. 17630v4 Announce Type: replace-cross Abstract: Circuit localization methods aim to identify the subset of model components responsible for specific behaviors in large language models, enabling detailed mechanistic analysis.

By Joakim Edin, Casper L. Christensen, R\'obert Csord\'as, Tuukka Ruotsalo, Zhengxuan Wu, Maria Maistro, Jing Huang, Lars Maal{\o}e