arXiv:2606. 17706v1 Announce Type: cross Abstract: Curriculum learning couples two design choices, how samples are scored by difficulty and how harder samples are paced into training, making it difficult to attribute observed gains to either component.
By Savini Kommalage, Sanka Mohottala, Asiri Gawesha, Dulara Madhusanka, Menan Velayuthan, Dharshana Kasthurirathna, Mahima Milinda Alwis Weerasinghe, Charith Abhayaratne
The paper investigates why curriculum learning—ordering training data from easy to hard—varies in effectiveness across reasoning tasks. By studying optimization dynamics, the authors introduce Relative Transfer, a measure of cross‑difficulty knowledge transfer, and use it to create Transfer‑aware Dynamic Curriculum Sampling (TDCS). Experiments show TDCS outperforms existing scheduling strategies on multiple reasoning benchmarks, offering a unified optimization‑based explanation for curriculum learning.
By Zhikai Ding, Ziyi Ye
arXiv:2608. 16164v1 Announce Type: new Abstract: Training locomotion policies for complex unstructured terrain requires a curriculum to avoid early exploration failures.
By Rocky Liu, Tengyu Liu, Baoxiong Jia, Fangwei Zhong, Xinyi Tong, Hongzhao Xie, Siyuan Huang
arXiv:2504. 07856v4 Announce Type: replace Abstract: Curriculum learning enhances Direct Preference Optimization (DPO) for aligning Large Language Models (LLMs), yet existing methods rely on a one-dimensional view of difficulty.
By Mengyang Li, Haozhan Geng, Zhong Zhang, Shuang Liu
arXiv:2603. 13761v2 Announce Type: replace Abstract: Curriculum learning--ordering training examples in a sequence to aid machine learning--takes inspiration from human learning, but has not gained widespread acceptance.
By Amogh Inamdar, Zhenwei Tang, Ashton Anderson, Richard Zemel
Consistency distillation has significantly accelerated the inference of diffusion models. In this work, we reveal an intriguing asymmetry: while Logit-Normal sampling priors are highly efficacious for standard iterative generation, consistency distillation exhibits a distinctly different difficulty profile (e.
CanvasAnneal is a curriculum‑guided reinforcement learning framework designed to improve Diffusion Language Models (DLMs) on complex reasoning and tool‑use tasks. It starts training by injecting reasoning traces from a stronger teacher model into the diffusion canvas, then gradually reduces this guidance so the model learns to generate reasoning independently. Experiments on mathematical reasoning and tool‑use benchmarks show that CanvasAnneal outperforms standard diffusion RL methods such as diffu‑GRPO on tasks like MATH500, Countdown, and Tau2, and accelerates reward improvement, though the gains vary by task.
By Blake Olson, Yuhang Song, Emmett McQuinn, Yuan Shangguan
arXiv:2602. 07339v2 Announce Type: replace Abstract: Diffusion-based trajectory planners can model multi-modal driving behavior, but their iterative denoising process introduces a latency bottleneck for real-time closed-loop deployment.
By Ruturaj Reddy, Hrishav Bakul Barua, Junn Yong Loo, Thanh Thi Nguyen, Ganesh Krishnasamy
arXiv:2609.40055v1 Announce Type: cross
Abstract: On-policy distillation (OPD) provides dense supervision directly on student-generated trajectories, making it an effective post-training strategy for...
By Jiacheng Qiu, Yunsoo Kim, Ruichen Xu, Jian Luo, Petar M. Djuri\'c, Sima Mofakham
arXiv:2609.10518v1 Announce Type: new
Abstract: fMRI foundation models increasingly aggregate heterogeneous data across brain states, cohorts, and acquisition settings, yet pretraining domains are co...
By Junfeng Xia, Wenhao Ye, Junxiang Zhang, Jiayu Zuo, Mo Wang, Quanying Liu
DyMD introduces a Distribution Matching Distillation framework that adapts teacher supervision and critic fitting to preserve interaction dynamics in few-step video generation. By employing temporal affinity–conditioned re‑noise sampling and dynamics‑guided fake‑score tracking, DyMD balances motion recovery with visual quality. The method distills a 14B teacher into a 1.3B student that achieves significant gains on embodied‑video benchmarks and downstream action planning tasks.
By Haojun Xu, Jie Huang, Xin Lu, Mingchen Zhong, Zihao Fan, Linjiang Huang, Si Liu
Group Relative Policy Optimization (GRPO) is effective when the current policy already samples useful reasoning trajectories, but it stalls on hard prompts whose correct solution modes lie outside the student's on-policy support. We propose TREK (Teacher-Routed Exploration via Forward KL), a simple staged procedure that uses distillation not for imitation but for exploration support expansion.