arXiv AI By Yi Nian, Tiankai Yang, Yudi Zhang, Qi Pan, Zelong Xu, Shenzhe Zhu, Qingqing Luan, Yue Huang, Xiangliang Zhang, Yue Zhao

DOG-DPO:Dynamic Optimization in Geometry for Safety Alignment

Read the original on arXiv AI →

arXiv:2606. 07678v1 Announce Type: cross Abstract: Safety alignment for large language models relies on preference data, but current pipelines often train on large, redundant datasets.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 10

Suan: Rectifying Direct Preference Safety Alignment in Large Language Models

The paper introduces Suan, a new preference optimization algorithm designed to improve safety alignment in large language models. Suan operates directly at the gradient level, avoiding traditional variational derivations, which yields more interpretable and robust training dynamics. Experiments show that Suan outperforms existing methods, achieving superior safety alignment while maintaining response utility.

By Oleksandr Cherednichenko, Roman Klypa
arXiv AI
Aug 26

Preference Data Selection for Mitigating the Alignment Tax in Large Language Models

The paper introduces BALIGN, a balanced data selection strategy designed to reduce catastrophic forgetting—referred to as the alignment tax—in large language models during preference-based alignment. By analyzing preference optimization gradients, the authors identify three data-centric features that influence parameter drift: the reference model's log-probability margin, token length differences between chosen and rejected responses, and TF‑IDF similarity to general capability corpora. BALIGN aggregates these features into a composite risk score to filter out high-risk preference samples, thereby preserving foundational capabilities while maintaining alignment gains with minimal computational overhead.

By Minsu Kim, Jianxun Lian, Xing Xie, Steven Euijong Whang