arXiv Machine Learning By Hyunseok Paeng

The Injection Paradox: Brand-Level Suppression in Safety-Trained LLM Recommendations via RAG Context Injection

Read the original on arXiv Machine Learning →

arXiv:2606. 09204v1 Announce Type: new Abstract: We present a reproducible failure mode of safety training in RAG-based LLM recommendation -- the Injection Paradox -- in which prompt injections embedded in retrieved documents backfire against the attacker, suppressing the target brand below the injection-free baseline.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.