arXiv Machine Learning By Tony Salomone, Deep Gandhi, Ali Asaria

How Modular Is a Frontier Mixture-of-Experts? A Pre-registered Causal Test in Which Apparent Expert Modularity Mostly Dissolves

Read the original on arXiv Machine Learning →

arXiv:2606. 25092v1 Announce Type: new Abstract: Sparse Mixture-of-Experts (MoE) models route each token to a few of many experts, inviting the hypothesis that experts form functional modules tied to capabilities or languages.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jun 10

From Observation to Intervention: A Causal Audit of Expert Importance in Mixture-of-Experts Models

arXiv:2606. 10703v1 Announce Type: new Abstract: Interpretability methods routinely use population-level summary statistics over observed model behaviour to license claims about the effects of targeted interventions on specific computations; in Pearl's terms, they treat rung-1 associational evidence as if it supported rung-2 interventional conclusions, a move whose validity is rarely tested.

By Leonard Engmann, Christian Medeiros Adriano, Holger Giese
arXiv Machine Learning
Aug 27

When Does Context Routing Help? A Systematic Study of Multi-Modal Fusion in Time Series Forecasting

The paper investigates when auxiliary context can genuinely improve multi‑modal time series forecasting. It identifies two necessary dataset‑level conditions: the target must not be dominated by a last‑value shortcut (low autocorrelation) and the context must provide additional information beyond history (non‑zero conditional mutual information). Experiments on a large mixture‑of‑experts model and several fusion mechanisms show that only when both conditions hold does context routing yield a substantial reduction in mean‑squared error; otherwise its contribution collapses to a capacity floor.

By Ruizhe Zhou, Gaoyuan Du, Xiaoyang Liu, Haoqi Yao, Deepayan Chakrabarti, Jiating Lin, Yixuan Shen
arXiv Machine Learning
Jun 5

Pattern Selectivity is Not Task-Causal Structure: A Cross-Architecture Mechanistic Study of Composed-Task Circuits in 1B-Class Language Models

arXiv:2606. 05378v1 Announce Type: new Abstract: We test whether a single screen-and-ablate recipe -- identify attention-head circuits by task-pattern selectivity, then verify by causal ablation against a matched-random null -- produces consistent mechanistic claims across model families.

By Yongzhong Xu