PMOPD: Task Ordering, Cycling, and Parameter-Update Subspace Protection in Multi-Teacher On-Policy Distillation
Read the original on Hugging Face Trending Papers →PMOPD introduces a projection-based approach to multi-teacher on-policy distillation, addressing the capability seesaw problem by constructing subspace memories from task-specific parameter displacements and projecting gradients and optimizer updates to avoid cross-task interference. It also includes a lightweight conflict probe for task interaction analysis, a task ordering strategy, and a cycling mechanism to balance subspace estimation and task revisitation. Experiments on Code, Reason, and Math tasks demonstrate that PMOPD improves all evaluated capabilities, raising average scores by 2.54 points on Qwen2.5-7B and 2.09 points on Llama-3.1-8B.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.