arXiv AI

DNAlign: Dynamic Null-Space Safe Alignment for LLMs

DNAlign is a lightweight alignment framework that uses control‑theoretic optimization and null‑space projection to steer large language models toward safe behavior while preserving core knowledge and response quality. By treating the LLM as a dynamic system, it introduces controllable perturbations that are restricted to a harmful‑related subspace derived from neutral hidden states, and a value function trained on human preference data adaptively optimizes these control signals. Extensive evaluations across multiple LLM backbones show that DNAlign consistently reduces harmful outputs while maintaining fluency, coherence, and factual utility, outperforming prior alignment baselines without sacrificing generation diversity.

arXiv AI
Aug 24

CLEAR: Continuous Latent Adapter Routing for Utility-Preserving LLM Safety Alignment

CLEAR is a conditional safety adaptation framework that employs a lightweight hidden‑state gate to continuously control the activation strength of a safety low‑rank adapter. It aims to reduce harmful completions while preserving the performance of the frozen backbone on benign prompts. Experiments on safety and utility benchmarks, including HarmBench and GSM8K, show that CLEAR significantly lowers HarmBench ASR and improves utility compared to globally applied safety tuning methods such as SFT or standard LoRA.

By Chengxiao Wang, Enyi Jiang, Xiaojing Liao, Sanmi Koyejo
arXiv AI
Oct 1

Faithful Dual-constrained Erasure for Robust LLM Safety Alignment

The paper introduces FDCU, a dual‑constrained subspace projection framework designed to improve machine unlearning for large language models. FDCU limits parameter updates with a dual‑masking rule that preserves general knowledge via Fisher Information while preventing the activation of spurious suppressors through the Principle of Minimal Functional Intervention. Experiments show that FDCU achieves state‑of‑the‑art robustness against retraining attacks while maintaining near‑lossless general utility, thereby ensuring durable safety for LLMs.

By Jiaqing Li, Shide Zhou, Zhibo Zhang, Yuxi Li, Tianlong Yu, Kailong Wang
arXiv AI
Jul 7

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models

arXiv:2607. 02914v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated remarkable capabilities across diverse applications, yet ensuring their simultaneous safety, helpfulness, and trustworthiness remains a persistent challenge.

By Jiyang Guan, Yong Xie, Jun Chen, Jiexi Liu, Zipeng Ye, Defeng Li, Jiayu Shen, Jialing Tao, Hui Xue