Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation
arXiv:2607. 15434v1 Announce Type: cross Abstract: Multi-agent systems routinely place one AI agent in authority over another.
The paper investigates how deception affects multi‑agent deliberation, finding that the key factor is the proportion of deceivers rather than the total number of agents. Defection rates—instances where initially correct agents adopt incorrect conclusions—grow linearly with the deceiver proportion, and large language model agents are vulnerable even when deceivers are a minority. The study also shows that coordination among deceivers can reduce their effectiveness and that the specific models involved influence susceptibility.
arXiv:2607. 15434v1 Announce Type: cross Abstract: Multi-agent systems routinely place one AI agent in authority over another.
arXiv:2608.22444v1 Announce Type: new Abstract: The unit of AI safety evaluation is still the individual model, yet language-model agents are increasingly deployed in interacting populations that rea...
arXiv:2607. 26120v1 Announce Type: new Abstract: Large Language Models (LLMs)-powered multi-agent systems are increasingly deployed in mixed-motive environments, where agents operate under asymmetric information and strategic deception due to conflicting or hidden objectives.
arXiv:2504.00285v2 Announce Type: replace Abstract: Large Language Models (LLMs) are effective at deceiving when prompted to do so. Models that demonstrate better performance on reasoning tasks are a...
arXiv:2605. 17480v3 Announce Type: replace Abstract: Multi-agent systems extend large language models (LLMs) by decomposing tasks among specialized agents, but their distributed decision process creates new attack surfaces.
arXiv:2609.35928v1 Announce Type: cross Abstract: Multi-agent LLM systems increasingly mix models from several providers, yet exposing each agent's underlying model identity to its peers significantl...
arXiv:2604.09746v2 Announce Type: replace-cross Abstract: As large language models (LLMs) are increasingly deployed as autonomous agents, understanding how strategic behavior emerges in multi-agent e...
arXiv:2605. 28114v2 Announce Type: replace Abstract: Language-model agents are moving from single-user assistants into persistent networks that build trust and reputation with one another, and the same models increasingly control physically embodied robots as well as software.
arXiv:2607. 03695v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly deployed in interacting populations, raising the question of what such populations come to believe collectively.
arXiv:2608. 11624v1 Announce Type: cross Abstract: Persuasion is a core dynamic of natural language communication, shaping how large language models (LLMs) update beliefs, resolve disagreements, and reach decisions.
Persuasion is a core dynamic of natural language communication, shaping how large language models (LLMs) update beliefs, resolve disagreements, and reach decisions. As LLMs increasingly debate, advise, and think collaboratively with humans and each other, resistance to harmful persuasion becomes a core requirement for reliable behavior.
arXiv:2604. 08169v2 Announce Type: replace Abstract: Alignment in LLMs is more brittle than commonly assumed: misalignment can be induced by adversarial prompts, benign fine-tuning, emergent misalignment, and goal misgeneralization.