arXiv:2606. 03648v1 Announce Type: cross Abstract: Adapting foundation large language models to a user's task or preferred style through fine-tuning can result in compromising the model's safety.
By Krishnapriya Vishnubhotla, Hillary Dawkins, Isar Nejadgholi, Svetlana Kiritchenko
arXiv:2603. 07445v2 Announce Type: replace-cross Abstract: Large language models (LLMs) often require fine-tuning (FT) to perform well on downstream tasks, but FT can induce safety-alignment drift even when the training dataset contains only benign data.
By Guoli Wang, Haonan Shi, Tu Ouyang, An Wang
The paper introduces a chance-constrained approach to fine‑tune large language models (LLMs) that limits the proportion of safety examples whose performance degrades beyond a set threshold relative to a reference model. By replacing the discontinuous violation indicator with a differentiable majorization, the authors derive a tractable, conservative constraint and a closed‑form, constraint‑aware gradient update that focuses on examples near or above the degradation threshold. Experiments on harmful fine‑tuning across three tasks and models show that this tail‑aware method consistently outperforms existing safety‑preserving baselines, suggesting that safety preservation should be treated as a reliability‑constrained optimization problem rather than average‑risk regularization.
By Taha Entesari, Mahyar Fazlyab
arXiv:2605.01913v2 Announce Type: replace-cross
Abstract: Fine-tuning safety-aligned language models for downstream tasks often leads to substantial degradation of refusal behavior, making models vul...
By Sadia Asif, Mohammad Mohammadi Amiri
arXiv:2606. 08044v1 Announce Type: cross Abstract: Large Language Model (LLM) safety has often been evaluated at the behavior level, which provides limited evidence of internal robustness, as these evaluations target outputs rather than representation-level vulnerability under intervention.
By Enyi Jiang, Anders Gj{\o}lbye, Yibo Jacky Zhang, Sanmi Koyejo
arXiv:2607. 09697v1 Announce Type: new Abstract: Existing safety mechanisms for multimodal large language models (MLLMs) face a fundamental trade-off between safety and utility.
By Jiayi Li, Kun Zhan