arXiv AI By Eli Corn, Daphna Weinshall

VISTA: Validation-Informed Trajectory Adaptation via Self-Distillation

Read the original on arXiv AI →

VISTA is an online self‑distillation framework that enforces consistency along a deep learning model’s optimization trajectory. It uses a validation‑informed Marginal Coverage score to identify earlier model states—called expert anchors—that retain specialized competence over distinct data regions. By integrating a coverage‑weighted ensemble of these anchors during training, VISTA regularizes the loss landscape, preserves learned knowledge, and improves robustness and generalization while cutting storage overhead by 90%.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 10

DUA-D2C: Dynamic Uncertainty Aware Method for Overfitting Remediation in Deep Learning

The paper introduces DUA-D2C, a Dynamic Uncertainty-Aware Divide2Conquer method that improves overfitting remediation in deep learning. It refines the traditional Divide2Conquer approach by dynamically weighting subset models based on a composite score of accuracy and normalized prediction entropy, allowing the central model to learn more from generalizable and confident edge models. The authors provide theoretical justification, show reduced model variance, and demonstrate significant generalization gains across image, audio, and text benchmarks, even when combined with standard regularizers like Dropout.

By Md. Saiful Bari Siddiqui, Md Mohaiminul Islam, Md. Golam Rabiul Alam
arXiv AI
Jun 9

Post-Trained MoE Can Skip Half Experts via Self-Distillation

arXiv:2605. 18643v2 Announce Type: replace-cross Abstract: Mixture-of-Experts (MoE) scales language models efficiently through sparse expert activation, and its dynamic variant further reduces computation by adjusting the activated experts in an input-dependent manner.

By Xingtai Lv, Li Sheng, Kaiyan Zhang, Yichen You, Siyan Gao, Xueheng Luo, Yuxin Zuo, Yuchen Fan, Junlin Yang, Ganqu Cui, Bingning Wang, Fan Yang, Youbang Sun, Ning Ding, Bowen Zhou
arXiv Machine Learning
Oct 2

Activation-Conditioned Self-Distillation

Activation-Conditioned Self-Distillation (ACSD) is a new on‑policy self‑distillation method that uses a frozen copy of the base model to extract a steering vector by contrasting activations from self‑generated trajectories that reach verified correct answers with all other trajectories. The student learns from next‑token distributions on its own prefixes, without needing reference text or teacher parameter updates, and is used alone at inference. Across five models, ACSD achieves the highest mean accuracy on four mathematical benchmarks, with notable gains on DeepSeek‑R1‑0528‑Qwen3‑8B and LiveCodeBench v6 compared to the OPSD baseline.

By Zhexi Lu, Subhajit Chaudhury, Tejaswini Pedapati, Keerthiram Murugesan, Lei Yu