arXiv AI By Prakhar Gupta, Garv Shah, Donghua Zhang

Self-Mined Hardness for Safety Fine-Tuning

Read the original on arXiv AI →

arXiv:2605. 03226v2 Announce Type: replace-cross Abstract: Safety fine-tuning of language models typically requires a curated adversarial dataset.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.

arXiv Machine Learning
Aug 11

When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs

arXiv:2608. 08542v1 Announce Type: new Abstract: Model merging has become the default way to give an aligned language model new skills without retraining: a practitioner folds task vectors from math, code, or domain specialists into a safety-aligned base using task arithmetic, TIES, or DARE.

By Yu Ma, Hongli Shi, Jing Li, Xinran Xu, Weiwei Hou