arXiv AI

A Unifying Lens on Supervised Fine-Tuning Through Target Distribution Design

arXiv:2606. 11189v1 Announce Type: cross Abstract: Supervised fine-tuning (SFT) typically maximizes the likelihood of every token in a demonstrated trajectory.

Hugging Face Trending Papers
Jul 6

LP-SFT: Local-Preserving Supervised Fine-Tuning via Multimodal Entropy Structure

Supervised fine-tuning (SFT) is the standard approach for adapting pretrained language models to downstream domains, yet it often improves target-domain behavior at the cost of degrading pre-existing capabilities. Standard cross-entropy fine-tuning promotes only the observed label token and leaves unconstrained how probability mass is redistributed over other plausible alternatives, potentially distorting the rich local preference structure learned during pretraining.

arXiv AI
Sep 11

Which Tokens Should SFT Actually Learn? A Token-Trimming Perspective on Mathematical Reasoning

The paper introduces Trimmed Logit-Gap SFT (TrimSFT), a token-level reweighting strategy that adjusts supervised fine-tuning loss based on the logit gap between the correct token and its strongest competitor. TrimSFT trims supervision from tokens that are either already mastered (large logit gap) or poorly supported (small or negative logit gap), focusing learning on tokens with intermediate logit gaps. Experiments on six base models across five mathematical reasoning benchmarks show that TrimSFT consistently outperforms standard SFT, achieving the best average performance on five of six models and up to +26.9 points on MATH500.

By Yaning Jia, Chunhui Zhang, Wenxuan Xu, Xingjian Diao, Xiaoyuan Wang, Soroush Vosoughi
arXiv AI
2d ago

PG-SFT: Balancing Capability Acquisition and Retention in Offline Agent Fine-Tuning

PG-SFT: Balancing Capability Acquisition and Retention in Offline Agent Fine-Tuning explores how to maintain a model’s existing abilities while teaching it new ones through supervised fine‑tuning on offline agent trajectories. The authors compare standard SFT, KL‑penalty, and update‑magnitude constraints, finding that these methods still degrade non‑target capabilities. They introduce Privilege‑Guided SFT (PG‑SFT), which uses turn‑level information gain to modulate supervision strength, achieving a better trade‑off between acquiring new skills and preserving existing ones, though with a slight drop in target‑task performance.

By Ronghua Li, Zi Liang, Zhishan Li, Shinan Liu
arXiv AI
Sep 10

Explainable Token-level Noise Filtering for LLM Fine-tuning Datasets

The paper introduces XTF, an explainable token‑level noise filtering framework for fine‑tuning large language models. XTF breaks down token contributions into reasoning importance, knowledge novelty, and task relevance, scores them, and masks gradients of noisy tokens to improve fine‑tuning. Experiments on math, code, and medicine tasks across seven LLMs show up to a 13.7% performance boost over standard fine‑tuning.

By Yuchen Yang, Wenze Lin, Enhao Huang, Zhixuan Chu, Hongbin Zhou, Lan Tao, Yiming Li, Zhan Qin, Kui Ren
arXiv Machine Learning
Aug 19

Data-DPO: Direct Preference Optimization for Target Model Data Selection in LLM Post-Training

Data-DPO is a target model‑oriented supervised fine‑tuning data selection method that uses one‑step probing of the target model to generate pairwise data preferences, trains a lightweight reward model to capture these preferences, and then selects a training subset by combining target‑model preference, external quality scores, and marginal diversity. Experiments on Vision‑Flan and LLaVA‑CoT demonstrate that Data‑DPO consistently outperforms existing data selection baselines across multiple data budgets and even surpasses full data training performance.

By Peng Sun, Yi Yang, Antong Zhang, Chunxiao Li, Yanbo Wang, Dianbo Liu, xin chen, Kai Yu, Lu Chen, Tianfan Fu