arXiv Computation and Language By Shaswat Patel, Vishvesh Trivedi, Yue Han, Yihuai Hong, Eunsol Choi

Bridging Latent Reasoning and Target-Language Generation via Retrieval-Transition Heads

Read the original on arXiv Computation and Language →

The paper investigates attention heads in multilingual Transformer models, distinguishing between retrieval heads that pull information from context and a newly identified class called Retrieval‑Transition heads (RTH) that direct the model toward a specific target language. Experiments across four multilingual benchmarks and two model families show that masking RTHs causes a larger performance drop than masking retrieval heads, indicating RTHs are crucial for chain‑of‑thought reasoning in multilingual LLMs. The study thus clarifies which attention heads are responsible for mapping to target languages, advancing our understanding of multilingual language models.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Computation and Language
Aug 27

Rethinking the Multilingual Reasoning Gap with Layer Swap

The study investigates the performance gap between native-language reasoning and English-pivoted reasoning in large language models. By creating extensive multilingual reasoning datasets and fine‑tuning specialists on Qwen/Qwen3-8B-Base, the authors find that the native reasoning gap is much smaller (1.9–3.5%) than previously reported. They analyze weight‑space changes, discover a language‑agnostic reasoning core in the middle layers, and propose a Layer Swap technique that transfers these mid‑layer updates from an English specialist to native specialists, effectively closing most of the gap while maintaining native chain‑of‑thought output.

By Maxence Lasbordes, Am\'elie Chatelain, Djam\'e Seddah