arXiv:2609.09099v1 Announce Type: new
Abstract: Curriculum learning is governed by several coupled design choices---how difficulty is defined, how examples are ordered, how much exposure each level r...
By Changho Shin, David Alvarez-Melis
The paper investigates why curriculum learning—ordering training data from easy to hard—varies in effectiveness across reasoning tasks. By studying optimization dynamics, the authors introduce Relative Transfer, a measure of cross‑difficulty knowledge transfer, and use it to create Transfer‑aware Dynamic Curriculum Sampling (TDCS). Experiments show TDCS outperforms existing scheduling strategies on multiple reasoning benchmarks, offering a unified optimization‑based explanation for curriculum learning.
By Zhikai Ding, Ziyi Ye
arXiv:2606. 00798v1 Announce Type: cross Abstract: Parameter compression of class-conditional diffusion models reveals an underexplored limitation in output-level distillation: the unconditional score branch remains unsupervised, leaving the classifier-free guidance gap underdetermined in the student.
By Abdullah Al Shafi, Kazi Saeed Alam, Sk Imran Hossain, Engelbert Mephu Nguifo
arXiv:2603. 13761v2 Announce Type: replace Abstract: Curriculum learning--ordering training examples in a sequence to aid machine learning--takes inspiration from human learning, but has not gained widespread acceptance.
By Amogh Inamdar, Zhenwei Tang, Ashton Anderson, Richard Zemel
CA-OPD is a confidence‑aware on‑policy distillation framework that improves structured visual prediction by using teacher confidence to selectively correct unreliable student transitions and gradually transfer rollout control to the student. The method aligns supervision with intervention decisions, providing direct cross‑entropy loss for corrected tokens and full predictive distribution for retained tokens. In a multi‑teacher setting for GUI grounding and OCR, CA‑OPD significantly outperforms the Qwen3.5‑0.8B baseline, achieving large gains on benchmarks such as ScreenSpot‑Pro and OCRBench‑v2 English.
By Menghao Li, Linjie Mu, Yin Wang, Haotian Hu, Yannian Gu, Lujiayi Xue, Fanyi Wang
arXiv:2402. 14035v4 Announce Type: replace-cross Abstract: Knowledge distillation from foundation models to compact domain models is challenging due to substantial gaps in capacity, architecture, and modality.
By Zichang Liu, Qingyun Liu, Yuening Li, Liang Liu, Anshumali Shrivastava, Shuchao Bi, Lichan Hong, Ed H. Chi, Zhe Zhao