arXiv Machine Learning

PRISM: Preference-Aware Influence Function Based Data Selection Method for Efficient Fine-Tuning

arXiv:2605. 21422v3 Announce Type: replace Abstract: As LLMs continue to scale up, improving training efficiency heavily relies on effective data utilization.

arXiv Machine Learning
Aug 19

Data-DPO: Direct Preference Optimization for Target Model Data Selection in LLM Post-Training

Data-DPO is a target model‑oriented supervised fine‑tuning data selection method that uses one‑step probing of the target model to generate pairwise data preferences, trains a lightweight reward model to capture these preferences, and then selects a training subset by combining target‑model preference, external quality scores, and marginal diversity. Experiments on Vision‑Flan and LLaVA‑CoT demonstrate that Data‑DPO consistently outperforms existing data selection baselines across multiple data budgets and even surpasses full data training performance.

By Peng Sun, Yi Yang, Antong Zhang, Chunxiao Li, Yanbo Wang, Dianbo Liu, xin chen, Kai Yu, Lu Chen, Tianfan Fu
arXiv AI
Sep 24

Fine-Tune, Then Rectify

The paper proposes a two‑stage framework that first fine‑tunes a large language model (LLM) and then rectifies its outputs, allocating limited labeled data optimally between the stages. It argues that the usual mean‑squared‑error objective for fine‑tuning misaligns with the downstream rectification, and instead suggests minimizing prediction‑error variance for mean estimation or a scalarized variance metric for general M‑estimation. Empirical results confirm that this variance‑based fine‑tuning, combined with optimal data allocation, yields significant efficiency gains over using either fine‑tuning or rectification alone, or using the conventional objective.

By Zikun Ye, Jinglong Zhao, Lei Wang
arXiv Computation and Language
Sep 24

RapidUn: Influence-Driven Parameter Reweighting for Efficient Large Language Model Unlearning

RapidUn is a parameter reweighting framework that uses influence estimates to guide LoRA-only updates for efficient unlearning of targeted behaviors in large language models. It operates in a practical PEFT setting with a small forget set and limited retain buffer, converting cross-sample influence into fixed sample-specific weights for weighted LoRA unlearning. Experiments on Llama‑3‑8B with Dolly‑15k and Alpaca‑57k datasets show RapidUn achieves lower trigger ASR than Fisher, GA, and LoReUn while preserving clean utility, and delivers a 77× wall‑clock speedup over clean‑corpus LoRA retraining, with additional evaluations supporting its effectiveness.

By Guoshenghui Zhao, Huawei Lin, Weijie Zhao