arXiv AI By Heming Zou, Yixiu Mao, Yun Qu, Qi Wang, Xiangyang Ji

Utility-Diversity Aware Online Batch Selection for LLM Supervised Fine-tuning

Read the original on arXiv AI →

arXiv:2510. 16882v4 Announce Type: replace-cross Abstract: Supervised fine-tuning (SFT) is a commonly used technique to adapt large language models (LLMs) to downstream tasks.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.

arXiv Machine Learning
22h ago

Data-DPO: Direct Preference Optimization for Target Model Data Selection in LLM Post-Training

arXiv:2608. 16926v1 Announce Type: new Abstract: Data selection in supervised fine-tuning aims to select a small set of effective samples from large-scale candidate data, reducing training cost while preserving model performance.

By Peng Sun, Yi Yang, Antong Zhang, Chunxiao Li, Yanbo Wang, Dianbo Liu, xin chen, Kai Yu, Lu Chen, Tianfan Fu