The paper introduces a chance-constrained approach to fine‑tune large language models (LLMs) that limits the proportion of safety examples whose performance degrades beyond a set threshold relative to a reference model. By replacing the discontinuous violation indicator with a differentiable majorization, the authors derive a tractable, conservative constraint and a closed‑form, constraint‑aware gradient update that focuses on examples near or above the degradation threshold. Experiments on harmful fine‑tuning across three tasks and models show that this tail‑aware method consistently outperforms existing safety‑preserving baselines, suggesting that safety preservation should be treated as a reliability‑constrained optimization problem rather than average‑risk regularization.
By Taha Entesari, Mahyar Fazlyab
arXiv:2607. 02781v1 Announce Type: cross Abstract: Inference-time alignment steers a frozen language model during decoding using auxiliary reward signals, avoiding the cost of repeated weight updates.
By Yaswanth Chittepu, Ativ Joshi, Sohini Chintala, Scott Niekum
arXiv:2603. 07445v2 Announce Type: replace-cross Abstract: Large language models (LLMs) often require fine-tuning (FT) to perform well on downstream tasks, but FT can induce safety-alignment drift even when the training dataset contains only benign data.
By Guoli Wang, Haonan Shi, Tu Ouyang, An Wang
arXiv:2608. 16068v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed as agents that rely on system prompts to use tools and complete tasks.
By Victor Ye Dong, Reid Pryzant, Yi Liu, Jian Jiao
arXiv:2603. 03305v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly used to generate executable outputs, JSON objects, and API calls, where a single syntax error can make the output unusable.
By Avinash Reddy, Thayne T. Walker, James S. Ide, Amrit Singh Bedi
arXiv:2606. 02530v1 Announce Type: new Abstract: Aligning Large Language Models (LLMs) with human values often degrades their general capabilities, termed the alignment tax.
By Hao Li, Jingkun An, Zijun Song, Pengyu Zhu, Rui Li, Hao Wang, Wendi Feng, Yesheng Liu, Lijun Li, Jin-Ge Yao, Lei Sha