Hugging Face Trending Papers

Safety Targeted Embedding Exploit via Refinement

Read the original on Hugging Face Trending Papers →

Safety training for large language models (LLMs) is conducted predominantly in English, leaving uncertain how well safety mechanisms generalize to low-resource languages and mixed-language code-switching. We show that this creates an epistemic gap in which models confidently generate harmful responses for inputs that fall outside the distribution of their safety training.

Summary generated by The Flow from the publisher's feed. The full article lives at Hugging Face Trending Papers.