arXiv:2606. 09396v1 Announce Type: cross Abstract: Supervised fine-tuning (SFT) is an efficient approach for downstream task adaptation and often serves as the initialization stage for reinforcement learning (RL), but it can show weaker generalization than RL.
By Ke Wang, Shuangqi Li, Mathieu Salzmann, Pascal Frossard
arXiv:2607. 04733v1 Announce Type: cross Abstract: Supervised fine-tuning (SFT) is the standard approach for adapting pretrained language models to downstream domains, yet it often improves target-domain behavior at the cost of degrading pre-existing capabilities.
By Yueyang Wang, Baolong Bi, Shuo Lu, Jingyuan Zhang
Supervised fine-tuning (SFT) is the standard approach for adapting pretrained language models to downstream domains, yet it often improves target-domain behavior at the cost of degrading pre-existing capabilities. Standard cross-entropy fine-tuning promotes only the observed label token and leaves unconstrained how probability mass is redistributed over other plausible alternatives, potentially distorting the rich local preference structure learned during pretraining.
arXiv:2606. 07527v1 Announce Type: cross Abstract: The prevailing paradigm for training LLMs has evolved to rely on a massive post-training phase consisting of SFT and RL.
By Michael Hassid, Yossi Adi, Roy Schwartz
The paper introduces Trimmed Logit-Gap SFT (TrimSFT), a token-level reweighting strategy that adjusts supervised fine-tuning loss based on the logit gap between the correct token and its strongest competitor. TrimSFT trims supervision from tokens that are either already mastered (large logit gap) or poorly supported (small or negative logit gap), focusing learning on tokens with intermediate logit gaps. Experiments on six base models across five mathematical reasoning benchmarks show that TrimSFT consistently outperforms standard SFT, achieving the best average performance on five of six models and up to +26.9 points on MATH500.
By Yaning Jia, Chunhui Zhang, Wenxuan Xu, Xingjian Diao, Xiaoyuan Wang, Soroush Vosoughi
PG-SFT: Balancing Capability Acquisition and Retention in Offline Agent Fine-Tuning explores how to maintain a model’s existing abilities while teaching it new ones through supervised fine‑tuning on offline agent trajectories. The authors compare standard SFT, KL‑penalty, and update‑magnitude constraints, finding that these methods still degrade non‑target capabilities. They introduce Privilege‑Guided SFT (PG‑SFT), which uses turn‑level information gain to modulate supervision strength, achieving a better trade‑off between acquiring new skills and preserving existing ones, though with a slight drop in target‑task performance.
By Ronghua Li, Zi Liang, Zhishan Li, Shinan Liu
arXiv:2606. 09856v1 Announce Type: cross Abstract: Post-training Large Language Models (LLMs) for reasoning typically focuses on deductive tasks such as mathematics and coding where correctness is verifiable.
By Liyi Zhang, Akshay K. Jagadish, Brenden M. Lake, Thomas L. Griffiths
arXiv:2604. 26170v2 Announce Type: replace Abstract: Adapting large language models (LLMs) to a targeted task efficiently and effectively remains a fundamental challenge.
By Ting-Wei Li, Sirui Chen, Jiaru Zou, Yingbing Huang, Tianxin Wei, Jingrui He, Hanghang Tong
The paper introduces XTF, an explainable token‑level noise filtering framework for fine‑tuning large language models. XTF breaks down token contributions into reasoning importance, knowledge novelty, and task relevance, scores them, and masks gradients of noisy tokens to improve fine‑tuning. Experiments on math, code, and medicine tasks across seven LLMs show up to a 13.7% performance boost over standard fine‑tuning.
By Yuchen Yang, Wenze Lin, Enhao Huang, Zhixuan Chu, Hongbin Zhou, Lan Tao, Yiming Li, Zhan Qin, Kui Ren
arXiv:2610.00814v1 Announce Type: cross
Abstract: Synthetic data are increasingly used to scale LLM training, yet more synthetic data do not necessarily produce better models. Useful synthetic data m...
By Yang Ba, Michelle V. Mancenido, Rong Pan
arXiv:2606. 18307v1 Announce Type: cross Abstract: Optimizing the training data distribution for Supervised Fine-Tuning (SFT) dictates the capability of Large Language Models (LLMs).
By Zefan Wang, Lincheng Li, Tianyu Yu, Yuan Yao
Data-DPO is a target model‑oriented supervised fine‑tuning data selection method that uses one‑step probing of the target model to generate pairwise data preferences, trains a lightweight reward model to capture these preferences, and then selects a training subset by combining target‑model preference, external quality scores, and marginal diversity. Experiments on Vision‑Flan and LLaVA‑CoT demonstrate that Data‑DPO consistently outperforms existing data selection baselines across multiple data budgets and even surpasses full data training performance.
By Peng Sun, Yi Yang, Antong Zhang, Chunxiao Li, Yanbo Wang, Dianbo Liu, xin chen, Kai Yu, Lu Chen, Tianfan Fu