arXiv:2609.39445v1 Announce Type: cross
Abstract: Time series foundation models (TSFMs) commonly adapt to new data by attaching a single trainable head to a frozen backbone, a one-size-fits-all setup...
By Hung Phan, Thuy T. Nguyen, Minh Ngoc Dinh, Nhat-Quang Tran
arXiv:2609.08115v1 Announce Type: new
Abstract: Mixture-of-Experts (MoE) pretraining relies on an auxiliary load-balancing loss (LBL) to drive per-expert utilization toward uniformity. Post-training...
By Jaedeok Lee, Keonwoo Kim, Dongyoon Han, Sangdoo Yun, Yera Choi, Haanju Yoo
The paper investigates self‑distillation techniques for language models by systematically varying three key design choices: the source of rollout tokens (student vs. teacher), the teacher coupling strategy (frozen or exponential moving average), and the KL divergence direction (reverse or forward). Experiments on Qwen2.5‑7B and Ministral‑3‑3B across 1,200 adaptation runs reveal that rollout source mainly affects acquisition on contradictory tasks, teacher coupling most strongly influences acquisition across all tasks, and KL direction impacts retention differently depending on the model. A controlled theoretical model reproduces these empirical trends, offering a unified framework for understanding acquisition‑retention trade‑offs in self‑distillation.
By Luis Zuin, Alexis Huet, Dario Rossi, Zied Ben Houidi
arXiv:2608. 11212v1 Announce Type: new Abstract: Top-k Mixture-of-Experts (MoE) routing is discontinuous, so a deployment-motivated numerical disturbance -- simulated 4-bit KV-cache quantization read by a protected BF16 gate -- pushes tokens across decision boundaries and flips which experts fire.
By Parvel Gu
The paper introduces component routing for self‑improving GUI agents, separating experience into locators, procedures, state facts, and lessons, and directing each to either the model weights or the prompt context. Experiments across three backbone families, two environments, and multiple seeds show that routing improves performance over whole‑trajectory baselines, with a rule based on recurrence and state‑conditionality accurately predicting the optimal destination. The study also analyzes how training dynamics and producer‑consumer differences affect the value of each destination, revealing that readout decreases for frequently recurring items when written to weights, while context gains grow with the information gap and weight gains shrink with the policy gap.
By Beining Wu, Zihao Ding, Jun Huang
arXiv:2608.20442v1 Announce Type: new
Abstract: Subliminal trait transfer allows a student model to acquire behavioral dispositions from teacher-generated data in which the trait is not semantically...
By Qinyang Xu