arXiv AI By Zezhong Wang, Xueyang Tang, Rui Lian, Yang Lou, Heqing Huang

Speculative Safety Honeypot: Toward Proactive Defense Against Multi-turn Agent Attacks

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

arXiv AI
Jun 16

Honeypot Protocol

arXiv:2604. 13301v1 Announce Type: cross Abstract: Trusted monitoring, the standard defense in AI control, is vulnerable to adaptive attacks, collusion, and strategic attack selection.

By Najmul Hasan
arXiv AI
Aug 6

Temporal Context Awareness: A Defense Framework Against Multi-turn Manipulation Attacks on Large Language Models

arXiv:2503. 15560v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly vulnerable to sophisticated multi-turn manipulation attacks, where adversaries strategically build context through seemingly benign conversational turns to circumvent safety measures and elicit harmful or unauthorized responses.

By Prashant Kulkarni, Assaf Namer