The paper investigates whether structural differences in circuits discovered by circuit discovery methods reflect distinct mechanisms. By varying input-token frequency while keeping the task constant, the authors find that although circuits appear specialized by frequency structurally, functional and representational analyses reveal no reliable differences, a phenomenon they call phantom specialization. Across multiple models and tasks, structurally distinct circuits implement the same computation, with core shared subgraphs recovering most of the performance and interchangeable internal representations confirmed by causal interventions.
By Alireza Bayat Makou, Jingcheng Niu, Subhabrata Dutta, Iryna Gurevych
The paper introduces Language Identity Head Ablation (LIHA), a causal method that zeroes individual attention heads in transformer models to measure language switch rates across multilingual prompts. Applying LIHA to GPT‑2 reveals a small set of first‑token broadcaster heads—most notably L6H1—that persistently attend to the initial prompt token and propagate language signals throughout generation, with compensatory head recruitment occurring hierarchically in higher layers. A controlled comparison between Qwen2.5‑1.5B‑Base and Qwen2.5‑1.5B‑Instruct shows that instruction tuning concentrates language‑identity influence in early layers, while experiments with Chinese and Russian confirm script‑specific first‑token broadcasting at layer 0.
By Arjun Pillai, Christian Hoang, Anjelo Jann Laroza
arXiv:2609.15545v1 Announce Type: new
Abstract: Hybrid language models can improve capability as well as efficiency, raising the question of how architectural complementarity becomes learned computat...
By Ke Cheng, Xin Xu, Yixiao Chen, Lei Xin, Jianbo Zhao, Fanhu Zeng, Yue Liu, Jun Zhang, Jie Jiang
arXiv:2603. 09161v2 Announce Type: replace-cross Abstract: Learning effective netlist representations is fundamentally constrained by the scarcity of labeled datasets, as real designs are protected by Intellectual Property (IP) and costly to annotate.
By Siyang Cai, Cangyuan Li, Haoyu Gao, Kun Wang, Yinhe Han, Ying Wang
arXiv:2606. 09899v1 Announce Type: cross Abstract: A central goal of mechanistic interpretability is to identify which internal components causally drive a language model's behavior.
By Luyang Zhang, Jialu Wang
arXiv:2610.00673v1 Announce Type: cross
Abstract: Looped language models increase effective depth by repeatedly applying a shared block of layers, but existing large-scale recipes require multi-stage...
By Andrei Marchenko, Viacheslav Bezrukov, Oleg Kashurin, Inessa Fedorova, Dmitry Bocharov, Yuliana Shakhvalieva, Maria Tikhonova, Valerii Ternovskii
arXiv:2606. 12818v1 Announce Type: cross Abstract: Irrelevant numbers in a prompt can shift language model judgments, producing anchoring effects in numerical reasoning.
By Hillary N. Owusu, Sarah Wiegreffe, Naomi H. Feldman
arXiv:2608. 05732v1 Announce Type: new Abstract: Controlling the behavior of large language models (LLMs) remains a critical challenge for AI alignment.
By Mehrshad Saadatinia, Parsa Razmara, Ardalan Aryashad, Ali Abbasi, Seyedarmin Azizi
arXiv:2511. 20849v2 Announce Type: replace-cross Abstract: We introduce a new tokenizer for language models that minimizes the average tokens per character, thereby reducing the number of tokens needed to represent text during training and to generate text during inference.
By Dong Dong, Weijie Su
arXiv:2609.37604v1 Announce Type: new
Abstract: Graph foundation models need a discrete token representation, but casting a graph as a generatable token sequence faces a structural obstacle: edges sp...
By Yuxiang Yao, Zijun Zhao
The study investigates how post‑training of large autoregressive language models (ARMs) into masked diffusion models (MDMs) affects their internal computation. Across two 7B ARM‑MDM families and four diagnostic tasks, the authors find that MDMs retain much of the ARM’s high‑attribution pathways on prefix‑dominant tasks, but reorganize computation toward earlier layers on globally constrained tasks. Component‑level probes reveal that ARMs depend on sharply specialized components, whereas MDMs show weaker specialization and more diffuse output‑space alignment.
By Injin Kong, Hyoungjoon Lee, Yohan Jo
arXiv:2609.24657v1 Announce Type: cross
Abstract: Language models can be adapted by changing the computations applied to individual tokens. Quantum circuits offer one such approach, but evaluating wi...
By Xiaoqiang Wang, Mengyang Xiong, Jun Dai, Bang Liu