arXiv Machine Learning By Peng Sun, Yi Yang, Antong Zhang, Chunxiao Li, Yanbo Wang, Dianbo Liu, xin chen, Kai Yu, Lu Chen, Tianfan Fu

Data-DPO: Direct Preference Optimization for Target Model Data Selection in LLM Post-Training

Read the original on arXiv Machine Learning →

Data-DPO is a target model‑oriented supervised fine‑tuning data selection method that uses one‑step probing of the target model to generate pairwise data preferences, trains a lightweight reward model to capture these preferences, and then selects a training subset by combining target‑model preference, external quality scores, and marginal diversity. Experiments on Vision‑Flan and LLaVA‑CoT demonstrate that Data‑DPO consistently outperforms existing data selection baselines across multiple data budgets and even surpasses full data training performance.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
2d ago

Scalable, Transferable Meta-network for Data Selection Requires a Different Loss (and Why the Obvious Choice is Problematic)

The paper introduces TESS, a scalable data‑selection framework that replaces per‑sample weights with a selection network to improve transferability across datasets and model sizes. It identifies instability in existing meta‑learning for training‑data selection (MTS) due to weight suppression and overreliance on easy features, and proposes a Pointwise Value Matching objective to address these issues. Experiments on large language model safety and instruction tuning show strong transfer from subsets to full corpora and from smaller to larger models.

By Zilin Du, Bowen Yang, Boyang Albert Li