arXiv AI By Cedric Caruzzo, Donggeun Yoo, Tae Soo Kim

Routing Divergence Is Not Evidence of Behavioral Influence in Same-Weight MoE Self-Distillation

Read the original on arXiv AI →

arXiv:2608. 15787v1 Announce Type: cross Abstract: Two Mixture-of-Experts (MoE) forward passes can share every weight yet route the same token through different experts.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
3d ago

Disentangling Self-Distillation: Measuring and Modeling Acquisition and Retention

The paper investigates self‑distillation techniques for language models by systematically varying three key design choices: the source of rollout tokens (student vs. teacher), the teacher coupling strategy (frozen or exponential moving average), and the KL divergence direction (reverse or forward). Experiments on Qwen2.5‑7B and Ministral‑3‑3B across 1,200 adaptation runs reveal that rollout source mainly affects acquisition on contradictory tasks, teacher coupling most strongly influences acquisition across all tasks, and KL direction impacts retention differently depending on the model. A controlled theoretical model reproduces these empirical trends, offering a unified framework for understanding acquisition‑retention trade‑offs in self‑distillation.

By Luis Zuin, Alexis Huet, Dario Rossi, Zied Ben Houidi
arXiv AI
2d ago

Not All Experience Belongs in the Weights: Component Routing for Self-Improving GUI Agents

The paper introduces component routing for self‑improving GUI agents, separating experience into locators, procedures, state facts, and lessons, and directing each to either the model weights or the prompt context. Experiments across three backbone families, two environments, and multiple seeds show that routing improves performance over whole‑trajectory baselines, with a rule based on recurrence and state‑conditionality accurately predicting the optimal destination. The study also analyzes how training dynamics and producer‑consumer differences affect the value of each destination, revealing that readout decreases for frequently recurring items when written to weights, while context gains grow with the information gap and weight gains shrink with the policy gap.

By Beining Wu, Zihao Ding, Jun Huang