The paper investigates attention heads in multilingual Transformer models, distinguishing between retrieval heads that pull information from context and a newly identified class called Retrieval‑Transition heads (RTH) that direct the model toward a specific target language. Experiments across four multilingual benchmarks and two model families show that masking RTHs causes a larger performance drop than masking retrieval heads, indicating RTHs are crucial for chain‑of‑thought reasoning in multilingual LLMs. The study thus clarifies which attention heads are responsible for mapping to target languages, advancing our understanding of multilingual language models.
By Shaswat Patel, Vishvesh Trivedi, Yue Han, Yihuai Hong, Eunsol Choi
arXiv:2608. 15964v1 Announce Type: cross Abstract: Language-specific competency (LSC) is the phenomenon of a language model performing better or worse depending on the language of the prompt.
By Ishika Agarwal, Arkajyoti Charaborty, Tanner Sorensen, Neha Gupta, Andreas Stolcke
The study investigates the performance gap between native-language reasoning and English-pivoted reasoning in large language models. By creating extensive multilingual reasoning datasets and fine‑tuning specialists on Qwen/Qwen3-8B-Base, the authors find that the native reasoning gap is much smaller (1.9–3.5%) than previously reported. They analyze weight‑space changes, discover a language‑agnostic reasoning core in the middle layers, and propose a Layer Swap technique that transfers these mid‑layer updates from an English specialist to native specialists, effectively closing most of the gap while maintaining native chain‑of‑thought output.
By Maxence Lasbordes, Am\'elie Chatelain, Djam\'e Seddah
arXiv:2609.37104v1 Announce Type: cross
Abstract: Open-source reasoning models provide unequal access to reasoning capability across languages. When a model can solve a problem but cannot deliver a c...
By Hongyang Li, Xiao Li, Caesar Wu, Gr\'egoire Danoy, Pascal Bouvry
arXiv:2607. 01002v1 Announce Type: cross Abstract: In long-context use, large language models frequently synthesize answers from the meaning of a relevant context span rather than literally copy-pasting them.
By Aryo Pradipta Gema, Beatrice Alex, Pasquale Minervini
arXiv:2606. 11198v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) systems inject external knowledge to improve LLM outputs, yet the format of injected content -- distinct from its semantic relevance -- can independently distort the model's attention distribution.
By Yuqi Zhang, Di Zhang