arXiv AI By Oleksandr Cherednichenko, Roman Klypa

Suan: Rectifying Direct Preference Safety Alignment in Large Language Models

Read the original on arXiv AI →

The paper introduces Suan, a new preference optimization algorithm designed to improve safety alignment in large language models. Suan operates directly at the gradient level, avoiding traditional variational derivations, which yields more interpretable and robust training dynamics. Experiments show that Suan outperforms existing methods, achieving superior safety alignment while maintaining response utility.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.