The paper introduces FDCU, a dual‑constrained subspace projection framework designed to improve machine unlearning for large language models. FDCU limits parameter updates with a dual‑masking rule that preserves general knowledge via Fisher Information while preventing the activation of spurious suppressors through the Principle of Minimal Functional Intervention. Experiments show that FDCU achieves state‑of‑the‑art robustness against retraining attacks while maintaining near‑lossless general utility, thereby ensuring durable safety for LLMs.
By Jiaqing Li, Shide Zhou, Zhibo Zhang, Yuxi Li, Tianlong Yu, Kailong Wang
arXiv:2603. 10938v2 Announce Type: replace-cross Abstract: Safe Reinforcement Learning from Human Feedback (RLHF) typically enforces safety through expected cost constraints, but the expectation captures only a single statistic of the cost distribution and fails to account for distributional uncertainty, particularly under heavy tails or rare catastrophic events.
By Yaswanth Chittepu, Ativ Joshi, Rajarshi Bhattacharjee, Scott Niekum
arXiv:2606. 15531v1 Announce Type: new Abstract: Fine-tuning aligned language models on benign tasks (e.
By Bohdan Turbal, Blossom Metevier, Max Springer, Aleksandra Korolova
arXiv:2608. 08383v1 Announce Type: cross Abstract: Steering vectors are a lightweight tool for controlling LLM behavior.
By Yuxiao Li, Gjergji Kasneci
The paper introduces a chance-constrained approach to fine‑tune large language models (LLMs) that limits the proportion of safety examples whose performance degrades beyond a set threshold relative to a reference model. By replacing the discontinuous violation indicator with a differentiable majorization, the authors derive a tractable, conservative constraint and a closed‑form, constraint‑aware gradient update that focuses on examples near or above the degradation threshold. Experiments on harmful fine‑tuning across three tasks and models show that this tail‑aware method consistently outperforms existing safety‑preserving baselines, suggesting that safety preservation should be treated as a reliability‑constrained optimization problem rather than average‑risk regularization.
By Taha Entesari, Mahyar Fazlyab
arXiv:2605.01913v2 Announce Type: replace-cross
Abstract: Fine-tuning safety-aligned language models for downstream tasks often leads to substantial degradation of refusal behavior, making models vul...
By Sadia Asif, Mohammad Mohammadi Amiri
arXiv:2608. 05045v1 Announce Type: cross Abstract: Released aligned large language models remain vulnerable to malicious downstream finetuning.
By Yuxuan Huang, Xingyu Zeng, Tianhang Zheng, Chaochao Lu
arXiv:2606. 00320v1 Announce Type: new Abstract: We present an online, distribution-free framework for controlling the Conditional Value-at-Risk (CVaR), extending conformal tail risk control to non-stationary and adversarial environments.
By Catherine Chen, Jingyan Shen, Zhun Deng, Lihua Lei
arXiv:2607. 23388v1 Announce Type: cross Abstract: As constrained learning becomes increasingly common, models are trained under explicit feasibility requirements to enforce fairness, safety, robustness, regulariza- tion, and physics or logic constraints.
By Xin Wang (Jeff), R. Tyrrell Rockafellar (Jeff), Xuegang (Jeff), Ban
NeuronGuard is a fine‑tuning defense for large language models that hardens them against both jailbreak and neuron‑level attacks. It redistributes safety signals across many neurons by identifying safety‑critical ones with per‑layer linear classifiers, enforcing refusal behavior when those neurons are ablated, and applying KL‑divergence regularization for consistency. A randomized gradient projection preserves task performance, and the authors provide a formal guarantee that NeuronGuard lowers the attack success rate upper bound, with experiments showing near‑zero success rates across multiple models and attack strategies.
By Anjun Gao, Yueyang Quan, Yufei Xia, Zhuqing Liu, Minghong Fang
arXiv:2606. 18697v1 Announce Type: new Abstract: Model-based learning agents use learned world models to predict future states, plan actions, and adapt to new environments.
By Yibin Hu, Xiaolin Sun, Zizhan Zheng
The paper introduces Suan, a new preference optimization algorithm designed to improve safety alignment in large language models. Suan operates directly at the gradient level, avoiding traditional variational derivations, which yields more interpretable and robust training dynamics. Experiments show that Suan outperforms existing methods, achieving superior safety alignment while maintaining response utility.
By Oleksandr Cherednichenko, Roman Klypa