arXiv AI By Yuxuan Huang, Xingyu Zeng, Tianhang Zheng, Chaochao Lu

Gradient Immunity: Null-Space Resistance to Malicious Fine-Tuning

Read the original on arXiv AI →

arXiv:2608. 05045v1 Announce Type: cross Abstract: Released aligned large language models remain vulnerable to malicious downstream finetuning.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Jul 2

Safety Targeted Embedding Exploit via Refinement

Safety training for large language models (LLMs) is conducted predominantly in English, leaving uncertain how well safety mechanisms generalize to low-resource languages and mixed-language code-switching. We show that this creates an epistemic gap in which models confidently generate harmful responses for inputs that fall outside the distribution of their safety training.