arXiv Computation and Language By Minjun Kim, Inho Won, Junghun Yuk, Dongyeon Kim, Jihyo Kim, KyungTae Lim

Distribution-aware Language Neuron Identification in Multilingual Large Language Models

Read the original on arXiv Computation and Language →

The paper introduces a new method for identifying language-specific neurons in multilingual large language models (mLLMs). Unlike previous entropy-based approaches that only consider positive activations, the proposed Distribution-aware Language Neuron selection uses pairwise overlap coefficients of full activation distributions, including negative values, to cluster languages. Experiments on two mLLMs and two held-out corpora show that this method isolates language-specific causal effects more effectively, achieving up to 4.9× higher on-target language damage per neuron while maintaining off-target language performance.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Computation and Language
Aug 28

Double Trouble: Bilingual Pretraining Leaves Language-Conditioned Effects in Shared-Language Representations

The study compares an English-only and a bilingual decoder-only model, each 310 M parameters, trained on eight diverse languages while controlling for English exposure, compute, and document overlap. After aligning on shared English vocabulary, the authors find that token embeddings appear similar, but the deeper hidden states used for prediction diverge across models. This hidden‑state mismatch grows through middle transformer layers and persists despite controls, indicating that contextual processing differs between the models. "whyItMatters":"The findings show that embedding alignment can conceal significant internal representation differences, which is crucial for any downstream work that assumes aligned multilingual models are interchangeable."

By Anjishnu Mukherjee, Ziwei Zhu, Antonios Anastasopoulos