arXiv:2609.39445v1 Announce Type: cross
Abstract: Time series foundation models (TSFMs) commonly adapt to new data by attaching a single trainable head to a frozen backbone, a one-size-fits-all setup...
By Hung Phan, Thuy T. Nguyen, Minh Ngoc Dinh, Nhat-Quang Tran
arXiv:2609.08115v1 Announce Type: new
Abstract: Mixture-of-Experts (MoE) pretraining relies on an auxiliary load-balancing loss (LBL) to drive per-expert utilization toward uniformity. Post-training...
By Jaedeok Lee, Keonwoo Kim, Dongyoon Han, Sangdoo Yun, Yera Choi, Haanju Yoo
The paper investigates self‑distillation techniques for language models by systematically varying three key design choices: the source of rollout tokens (student vs. teacher), the teacher coupling strategy (frozen or exponential moving average), and the KL divergence direction (reverse or forward). Experiments on Qwen2.5‑7B and Ministral‑3‑3B across 1,200 adaptation runs reveal that rollout source mainly affects acquisition on contradictory tasks, teacher coupling most strongly influences acquisition across all tasks, and KL direction impacts retention differently depending on the model. A controlled theoretical model reproduces these empirical trends, offering a unified framework for understanding acquisition‑retention trade‑offs in self‑distillation.
By Luis Zuin, Alexis Huet, Dario Rossi, Zied Ben Houidi
arXiv:2608. 11212v1 Announce Type: new Abstract: Top-k Mixture-of-Experts (MoE) routing is discontinuous, so a deployment-motivated numerical disturbance -- simulated 4-bit KV-cache quantization read by a protected BF16 gate -- pushes tokens across decision boundaries and flips which experts fire.
By Parvel Gu
The paper introduces component routing for self‑improving GUI agents, separating experience into locators, procedures, state facts, and lessons, and directing each to either the model weights or the prompt context. Experiments across three backbone families, two environments, and multiple seeds show that routing improves performance over whole‑trajectory baselines, with a rule based on recurrence and state‑conditionality accurately predicting the optimal destination. The study also analyzes how training dynamics and producer‑consumer differences affect the value of each destination, revealing that readout decreases for frequently recurring items when written to weights, while context gains grow with the information gap and weight gains shrink with the policy gap.
By Beining Wu, Zihao Ding, Jun Huang
arXiv:2608.20442v1 Announce Type: new
Abstract: Subliminal trait transfer allows a student model to acquire behavioral dispositions from teacher-generated data in which the trait is not semantically...
By Qinyang Xu
arXiv:2608. 12957v1 Announce Type: new Abstract: Group Relative Policy Optimization (GRPO) learns from reward differences within a rollout group, but receives no useful relative signal when every sampled response is incorrect.
By Yubo Zhang, Xinhong Ma, Zezhong Tan, Ziqiang Dong
arXiv:2608. 09826v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards yields no group-relative signal when rollout groups are uniformly correct or uniformly wrong, which account for 63.
By Yubo Jiang, Fengying Xie, Zhiguo Jiang, Haopeng Zhang
The paper investigates how on-policy self‑distillation can alter a model’s behavior by conditioning on privileged information. It contrasts attractive self‑distillation, which pulls a model toward a privileged teacher, with repulsive self‑distillation, which pushes it away, showing that attraction reduces exploratory reasoning while repulsion lengthens responses and can destabilize the model. The authors propose a contrastive self‑distillation objective that combines attraction to a correct‑solution teacher with repulsion from an incorrect‑solution teacher, finding that this approach improves reasoning performance across various model types while keeping response lengths stable.
By Anton Baumann, Akmal Ashirmatov, Leo Schmidt-Traub, Frederike L\"ubeck, Jonas H\"ubotter, Thomas Kleine Buening, Andreas Krause
The paper investigates subliminal learning, where hidden traits from a teacher model are transferred to a student during distillation. It introduces trait‑direction drift as the underlying mechanism, showing that biased generation creates measurable preference gaps that accumulate into behavioral transfer during fine‑tuning. The authors propose probe‑space corridor regularization, a targeted defense that constrains drift along a calibrated trait direction, significantly reducing hidden‑trait transfer while maintaining task performance.
By Zhixuan Liu, Zhichen Dong, Yuyu Fan, Xiangtian Li, Chao Yang
arXiv:2607. 21692v2 Announce Type: replace Abstract: Sparse attention prunes a long context to the blocks a model needs, and the usual selector is distilled from a dense teacher's attention.
By Jim Allchin
The paper investigates how the four‑stream manifold‑constrained hyper‑connection (mHC) residual pathway in DeepSeek‑V4‑Flash is actually used. It finds that read/write routing is concentrated, typically involving only two streams per block, and that the dominant stream shifts across layers while representations stay directionally distinct. Residual mixing is modest, mainly in early layers, and late mixing contributes little to performance, whereas early mixing is crucial for perplexity and task scores.
By Pengxiang Zhao, Xing Li, Xianzhi Yu, Wei Guo, Zhenhua Dong