arXiv AI By Yanchen Yin, Dongqi Han, Linghui Li

Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models

Read the original on arXiv AI →

arXiv:2606. 28153v1 Announce Type: cross Abstract: Jailbreak attacks bypass LLM safety alignment, yet their mechanisms remain poorly understood.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.