arXiv AI By Jinhan Li, Kexian Tang, Yihan Xu, Zhuorui Ye, Kaifeng Lyu

Beyond Safe Data: Pretraining-Stage Alignment with Regular Safety Reflection

Read the original on arXiv AI →

arXiv:2606. 19168v1 Announce Type: new Abstract: To achieve deeper safety alignment for large language models (LLMs), recent efforts have studied how to push safety interventions earlier into the pretraining stage, primarily by filtering unsafe data or rewriting it into safer forms.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.