arXiv AI By Wei Duan, Li Qian

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message

Read the original on arXiv AI →

arXiv:2507. 04673v2 Announce Type: replace Abstract: The rise of conversational interfaces has greatly enhanced LLM usability by leveraging dialogue history for sophisticated reasoning.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.

arXiv AI
Aug 6

Temporal Context Awareness: A Defense Framework Against Multi-turn Manipulation Attacks on Large Language Models

arXiv:2503. 15560v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly vulnerable to sophisticated multi-turn manipulation attacks, where adversaries strategically build context through seemingly benign conversational turns to circumvent safety measures and elicit harmful or unauthorized responses.

By Prashant Kulkarni, Assaf Namer