arXiv:2608. 14320v1 Announce Type: new Abstract: The anchoring effect is a cognitive bias in which an initial reference value shifts a later judgment toward itself.
By Yiderigun Borjigin, Alexander Hermann, Christian Cyron, Roland Aydin
arXiv:2609.25602v1 Announce Type: new
Abstract: In language models, the choice between believing the prompt and believing the weights is made by a handful of identifiable attention heads. Instruction...
By Shubham Santosh Pandere, Gautam Ranka, Ritika Varshney, Navya Deshmukh, Roushni Sareen, Roshan Kumar Singh
The paper introduces FTB Graph, a method for mapping the causal circuitry that determines the first-token language identity in multilingual language models. Using Edge Attribution Patching and exact activation patching across six architectures (GPT‑2, BLOOM‑560M, Pythia‑1B/2.8B, Qwen2.5‑1.5B Base/Instruct), the authors extract directed acyclic graphs that reveal deep or mid‑to‑deep broadcasting hubs, with notable differences among models. The study finds that first‑token routing is largely established during pretraining and largely preserved by instruction tuning, while linear gradient approximations can diverge from causal interventions, underscoring the need for exact‑patching verification.
By Arjun Pillai, Christian Hoang, Anjelo Laroza
arXiv:2604. 04385v5 Announce Type: replace-cross Abstract: We localize the policy routing mechanism in alignment-trained language models.
By Gregory N. Frank
The paper introduces a three-level evaluation framework—behavioral deployment, LM-head readout, and probe recoverability—to distinguish whether a language model fails a syntactic test by not encoding structure or by failing to use it. Using a trilingual control-dependency benchmark, the authors find that probe recoverability consistently exceeds LM-head readout, which in turn exceeds behavioral deployment across seven models and three languages, with the largest gap observed in Qwen3-0.6B Instruct. Layer-localized activation patching shows that instruction tuning shifts the decoded layer later, suggesting decoding favors surface shortcuts and that behavioral evaluation understates what models encode while probing alone overstates what they deploy.
By Zhenyan Lu, He Wang, Xiaohui Huang
arXiv:2606. 01202v1 Announce Type: new Abstract: Language models do not simply choose an answer at the output layer.
By Shailesh Rana