Hugging Face Trending Papers

PMOPD: Task Ordering, Cycling, and Parameter-Update Subspace Protection in Multi-Teacher On-Policy Distillation

Read the original on Hugging Face Trending Papers →

PMOPD introduces a projection-based approach to multi-teacher on-policy distillation, addressing the capability seesaw problem by constructing subspace memories from task-specific parameter displacements and projecting gradients and optimizer updates to avoid cross-task interference. It also includes a lightweight conflict probe for task interaction analysis, a task ordering strategy, and a cycling mechanism to balance subspace estimation and task revisitation. Experiments on Code, Reason, and Math tasks demonstrate that PMOPD improves all evaluated capabilities, raising average scores by 2.54 points on Qwen2.5-7B and 2.09 points on Llama-3.1-8B.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv Machine Learning
2d ago

No Task Vector Is an Island: A Comprehensive Study on the Composability of Task Vectors from On-Policy Distillation

arXiv:2609.39405v1 Announce Type: new Abstract: Task vectors provide a simple mechanism for composing learned capabilities through model merging. However, the composability of task vectors produced b...

By Jingang Zhou, Feiyu Han, Han Zhu, Yuyi Zhou, Ruiyang Zhang, Jian Xu, Sirui Gao, Qingpei Guo, Xu-Yao Zhang
arXiv Machine Learning
Jul 7

Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective Recipe

arXiv:2605. 03677v2 Announce Type: replace Abstract: On-policy distillation (OPD) has recently emerged as an effective post-training paradigm for consolidating the capabilities of specialized expert models into a single student model.

By Wenjin Hou, Shangpin Peng, Weinong Wang, Zheng Ruan, Yue Zhang, Zhenglin Zhou, Mingqi Gao, Yifei Chen, Kaiqi Wang, Hongming Yang, Chengquan Zhang, Zhuotao Tian, Han Hu, Yi Yang, Fei Wu, Hehe Fan