Data-DPO is a target model‑oriented supervised fine‑tuning data selection method that uses one‑step probing of the target model to generate pairwise data preferences, trains a lightweight reward model to capture these preferences, and then selects a training subset by combining target‑model preference, external quality scores, and marginal diversity. Experiments on Vision‑Flan and LLaVA‑CoT demonstrate that Data‑DPO consistently outperforms existing data selection baselines across multiple data budgets and even surpasses full data training performance.
By Peng Sun, Yi Yang, Antong Zhang, Chunxiao Li, Yanbo Wang, Dianbo Liu, xin chen, Kai Yu, Lu Chen, Tianfan Fu
arXiv:2501. 12147v2 Announce Type: replace-cross Abstract: Selecting appropriate training data is crucial for instruction fine-tuning of large language models (LLMs), which aims to (1) elicit strong capabilities, and (2) achieve balanced performance across different tasks.
By Qirun Dai, Dylan Zhang, Jiaqi W. Ma, Hao Peng
arXiv:2606. 18307v1 Announce Type: cross Abstract: Optimizing the training data distribution for Supervised Fine-Tuning (SFT) dictates the capability of Large Language Models (LLMs).
By Zefan Wang, Lincheng Li, Tianyu Yu, Yuan Yao
arXiv:2606. 09396v1 Announce Type: cross Abstract: Supervised fine-tuning (SFT) is an efficient approach for downstream task adaptation and often serves as the initialization stage for reinforcement learning (RL), but it can show weaker generalization than RL.
By Ke Wang, Shuangqi Li, Mathieu Salzmann, Pascal Frossard
arXiv:2509. 03647v2 Announce Type: replace-cross Abstract: Large language models (LLMs) increasingly serve as automated evaluators, yet they suffer from "self-preference bias": a tendency to favor their own outputs over those of other models.
By Dani Roytburg, Matthew Bozoukov, Matthew Nguyen, Jou Barzdukas, Simon Fu, Narmeen Oozeer
arXiv:2609.36659v1 Announce Type: new
Abstract: The strong generalization performance of on-policy post-training paradigms has motivated studies of their parameter update behaviors. However, these st...
By Shufan Shen, Zhongni Hou, Junshu Sun, Yufei Zhang, Wei Lin, Guojun Yin, Qingming Huang, Shuhui Wang