arXiv AI By Yuxuan Huang, Xingyu Zeng, Tianhang Zheng, Chaochao Lu

Gradient Immunity: Null-Space Resistance to Malicious Fine-Tuning

Read the original on arXiv AI →

arXiv:2608. 05045v1 Announce Type: cross Abstract: Released aligned large language models remain vulnerable to malicious downstream finetuning.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.

Hugging Face Trending Papers
Jul 2

Safety Targeted Embedding Exploit via Refinement

Safety training for large language models (LLMs) is conducted predominantly in English, leaving uncertain how well safety mechanisms generalize to low-resource languages and mixed-language code-switching. We show that this creates an epistemic gap in which models confidently generate harmful responses for inputs that fall outside the distribution of their safety training.