Teacher–student curriculum learning
Read the original on OpenAI Blog →The Flow has not summarised this story yet — read it at OpenAI Blog.
The Flow has not summarised this story yet — read it at OpenAI Blog.
On-policy distillation (OPD) transfers teacher capabilities by supervising trajectories sampled from the student's own policy, yet its generalization behavior remains poorly understood, as most studies evaluate OPD on a single domain and on benchmarks close to the training data. We present a controlled study that varies one generalization factor at a time, from in-domain distribution shifts to cross-domain transfer and the multi-teacher setting.
arXiv:2610.08778v1 Announce Type: new Abstract: Large language models (LLMs) have become increasingly capable problem solvers, but being able to solve a problem is not the same as being able to teach...
arXiv:2609.23088v1 Announce Type: new Abstract: Educational foundation models must solve problems, understand curriculum structure, diagnose learner difficulties, and provide appropriate instructiona...
arXiv:2607. 28647v1 Announce Type: cross Abstract: This paper presents ConnectED, a human-centered AI system that supports the full instructional lifecycle in Vietnamese education by linking curriculum-aligned lesson design, interactive student learning, and feedback-driven refinement.
PersonaPath is a new benchmark for knowledge‑centric personalized learning path planning, pairing 2,000 learner personas with a hierarchical knowledge graph of 347 textbooks, 1,751 units, and 4,092 concepts across 77 subjects. The study evaluates large language models on this benchmark, finding that even the best model achieves only a 29.5% final pass rate in Basic Education and fails to exceed 44.7% in tailoring paths to individual learners, highlighting a significant adaptivity gap. This work underscores the challenge of moving beyond exercise‑centric recommendation toward goal‑oriented, curriculum‑scale guidance.
arXiv:2606. 17706v1 Announce Type: cross Abstract: Curriculum learning couples two design choices, how samples are scored by difficulty and how harder samples are paced into training, making it difficult to attribute observed gains to either component.