arXiv AI By Sihui Dai, Mann Patel

What Do Safety-Aligned LLMs Learn From Mixed Compliance Demonstrations?

Read the original on arXiv AI →

arXiv:2606. 20508v1 Announce Type: new Abstract: Prior work has shown that in-context demonstrations can jailbreak language models, but it remains unclear how models interpret different types of compliance demonstrations.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.