arXiv Statistics ML

Raw-Routed Mixture of Adapters: A Causal Intervention for Routing Collapse in Time Series Foundation Models

Hugging Face Trending Papers
Sep 17

Should This Case Be Adapted? Prediction Fragmentation Controls Test-Time Adaptation

The paper investigates test‑time adaptation for medical image segmentation, showing that a fixed adaptation horizon can harm many individual cases. It introduces prediction fragmentation—a measure of disagreement between the source model and the adapted mask—to predict harmful adaptation without extra labels or backward passes. Using a case‑level router based on this metric, the authors reduce harmful adaptation on cardiac MRI from 58.7% to 20% while maintaining accuracy.

arXiv Computer Vision
Sep 18

Should This Case Be Adapted? Prediction Fragmentation Controls Test-Time Adaptation

The paper introduces a method for deciding whether to adapt a frozen segmentation model at test time, arguing that a fixed adaptation horizon conflates two distinct decisions: how far to adapt and whether to adapt at all. By measuring disagreement geometry—called prediction fragmentation—between the source model and the adapted mask, the authors predict harmful accepted area (HA) without extra labels or backward passes, achieving strong correlation across three medical benchmarks. A case‑level router built on this metric reduces HA significantly while maintaining or improving Dice scores, and the approach generalizes across architectures and domains.

By Lili Wang, Jing Li, Xiaowen Sun, Xiangyu Hu, Zhuangzhuang Gu, Jian Liu, Srihari Nelakuditi, Yan Tong
arXiv Machine Learning
Aug 27

When Does Context Routing Help? A Systematic Study of Multi-Modal Fusion in Time Series Forecasting

The paper investigates when auxiliary context can genuinely improve multi‑modal time series forecasting. It identifies two necessary dataset‑level conditions: the target must not be dominated by a last‑value shortcut (low autocorrelation) and the context must provide additional information beyond history (non‑zero conditional mutual information). Experiments on a large mixture‑of‑experts model and several fusion mechanisms show that only when both conditions hold does context routing yield a substantial reduction in mean‑squared error; otherwise its contribution collapses to a capacity floor.

By Ruizhe Zhou, Gaoyuan Du, Xiaoyang Liu, Haoqi Yao, Deepayan Chakrabarti, Jiating Lin, Yixuan Shen
arXiv Machine Learning
Jun 10

From Observation to Intervention: A Causal Audit of Expert Importance in Mixture-of-Experts Models

arXiv:2606. 10703v1 Announce Type: new Abstract: Interpretability methods routinely use population-level summary statistics over observed model behaviour to license claims about the effects of targeted interventions on specific computations; in Pearl's terms, they treat rung-1 associational evidence as if it supported rung-2 interventional conclusions, a move whose validity is rarely tested.

By Leonard Engmann, Christian Medeiros Adriano, Holger Giese
arXiv AI
1d ago

Signal-Routed Temperature Scaling: Low-Capacity Risk-Conditioned Calibration for Small Validation Budgets

Signal‑Routed Temperature Scaling (SRTS‑BCE) is a 10‑parameter, argmax‑preserving calibration method that separates calibration objectives from adaptive capacity. It cross‑fits a correctness‑risk score over six logit statistics and assigns a top‑label BCE temperature to each of three risk groups, generalizing TvA‑TS when K=1. Experiments on fine‑tuned CIFAR‑100 and ViT‑B/16 show that SRTS‑BCE reduces ECE from 1.65 to 0.96 with a small calibration budget, outperforming higher‑capacity SMART+BCE when only 250 examples are available, and revealing a budget‑dependent ranking reversal on Swin‑T.

By Wenhao Liang, Liangwei Nathan Zheng, Lin Yue, Wei Emma Zhang, Mingyu Guo, Olaf Maennel, Weitong Chen
arXiv Machine Learning
Aug 27

GRIP: Algorithm-Agnostic Machine Unlearning for Mixture-of-Experts via Geometric Router Constraints

The paper introduces GRIP, an algorithm‑agnostic framework for machine unlearning in Mixture‑of‑Experts large language models. GRIP enforces hard geometric constraints on router updates, projecting gradient changes into the null space of the retain set’s routing matrix to prevent routing manipulation. Two variants—training‑time stochastic projection and post‑training analytical correction—show significant improvements in routing stability, retain accuracy, and resistance to white‑box adversarial recovery across two MoE models.

By Andy Zhu, Rongzhe Wei, Yupu Gu, Pan Li