arXiv AI By Jisheng Dang, Yushuo Zhao, Dewei Liu, Junfeng Fang, Bimei Wang, Tiantian Rao, Hong Peng, Bin Hu, Tat-Seng Chua

DNAlign: Dynamic Null-Space Safe Alignment for LLMs

Read the original on arXiv AI →

DNAlign is a lightweight alignment framework that uses control‑theoretic optimization and null‑space projection to steer large language models toward safe behavior while preserving core knowledge and response quality. By treating the LLM as a dynamic system, it introduces controllable perturbations that are restricted to a harmful‑related subspace derived from neutral hidden states, and a value function trained on human preference data adaptively optimizes these control signals. Extensive evaluations across multiple LLM backbones show that DNAlign consistently reduces harmful outputs while maintaining fluency, coherence, and factual utility, outperforming prior alignment baselines without sacrificing generation diversity.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.