arXiv Machine Learning

How Modular Is a Frontier Mixture-of-Experts? A Pre-registered Causal Test in Which Apparent Expert Modularity Mostly Dissolves

arXiv:2606. 25092v1 Announce Type: new Abstract: Sparse Mixture-of-Experts (MoE) models route each token to a few of many experts, inviting the hypothesis that experts form functional modules tied to capabilities or languages.

arXiv Machine Learning
Jun 10

From Observation to Intervention: A Causal Audit of Expert Importance in Mixture-of-Experts Models

arXiv:2606. 10703v1 Announce Type: new Abstract: Interpretability methods routinely use population-level summary statistics over observed model behaviour to license claims about the effects of targeted interventions on specific computations; in Pearl's terms, they treat rung-1 associational evidence as if it supported rung-2 interventional conclusions, a move whose validity is rarely tested.

By Leonard Engmann, Christian Medeiros Adriano, Holger Giese
arXiv Machine Learning
Aug 27

When Does Context Routing Help? A Systematic Study of Multi-Modal Fusion in Time Series Forecasting

The paper investigates when auxiliary context can genuinely improve multi‑modal time series forecasting. It identifies two necessary dataset‑level conditions: the target must not be dominated by a last‑value shortcut (low autocorrelation) and the context must provide additional information beyond history (non‑zero conditional mutual information). Experiments on a large mixture‑of‑experts model and several fusion mechanisms show that only when both conditions hold does context routing yield a substantial reduction in mean‑squared error; otherwise its contribution collapses to a capacity floor.

By Ruizhe Zhou, Gaoyuan Du, Xiaoyang Liu, Haoqi Yao, Deepayan Chakrabarti, Jiating Lin, Yixuan Shen
arXiv Machine Learning
Jun 5

Pattern Selectivity is Not Task-Causal Structure: A Cross-Architecture Mechanistic Study of Composed-Task Circuits in 1B-Class Language Models

arXiv:2606. 05378v1 Announce Type: new Abstract: We test whether a single screen-and-ablate recipe -- identify attention-head circuits by task-pattern selectivity, then verify by causal ablation against a matched-random null -- produces consistent mechanistic claims across model families.

By Yongzhong Xu
arXiv Computation and Language
Sep 10

How Fragile Is Safety Alignment at Frontier Scale? A Single-Direction Attack on a 320B MoE

The paper demonstrates that a single-direction white‑box attack, which projects a ‘refusal direction’ from a language model’s weights, remains effective against a 320B‑parameter mixture‑of‑experts (MoE) model (GLM‑5.3‑Flash). The attack requires only a few hundred contrastive prompts and no gradient training, and it reduces refusal behavior by up to 89 percentage points across seven harmful‑content benchmarks while leaving overall capability unchanged. However, the attack’s impact is distributed across multiple components—attention, dense, and routed‑expert writers—so that only a joint intervention removes most of the refusal ability, and the conventional module‑name matching approach fails to capture this effect in MoE architectures.

By Yi Shi, Tanyu Chen, Kai Shen
arXiv Machine Learning
Sep 18

Intrinsic Sequence-Likelihood Confidence in Retrieval-Dominated Extractive QA: Two Pre-Specified Negatives, and What They Do and Do Not Attribute

The paper investigates whether confidence signals from fine‑tuned large language models can improve extractive question answering that relies heavily on retrieval. Experiments on four 7‑9B model families show that retrieval alone recovers 92–99.8% of the best possible accuracy, leaving little room for confidence‑based routing or adaptation to help. The sequence‑likelihood confidence metric, even after recalibration or temperature scaling, fails to provide a statistically significant benefit across different correctness criteria and answer lengths, and the study ultimately offers a set of pre‑specified negatives with explicit dependencies as its main contribution.

By Gunwoo Lee, Changmin Sung, Sang-Hwan Gwak, Ina Kim, Ji-Young Choi, Kyong-Ha Lee
arXiv Machine Learning
Aug 4

Capability Provenance in Language Models: A Case Study in Social Reasoning

arXiv:2606. 19625v2 Announce Type: replace-cross Abstract: We use training-data attribution as an interpretable tool for capability discovery, mapping which regions of the pretraining corpus support social-reasoning versus STEM-reasoning in OLMo3-7B.

By Glenn Matlin, Chandreyi Chakraborty, Saehee Eom, Mika Okamoto, Rayan Castilla, Louis Jaburi, Alvin Deng, Taywon Min, Lucia Quirke, Stella Biderman, Mark Riedl