arXiv AI By Aarush Sinha, Arion Das, Soumyadeep Nag, Charan Karnati, Shravani Nag, Chandra Vadhan Raj, Aman Chadha, Vinija Jain, Suranjana Trivedy, Amitava Das

CONSCIENTIA: Can LLM Agents Learn to Strategize? Emergent Deception and Trust in a Multi-Agent NYC Simulation

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

Hugging Face Trending Papers
Aug 12

Learning to Persuade Exposes How Easily LLMs Abandon Correct Beliefs

Persuasion is a core dynamic of natural language communication, shaping how large language models (LLMs) update beliefs, resolve disagreements, and reach decisions. As LLMs increasingly debate, advise, and think collaboratively with humans and each other, resistance to harmful persuasion becomes a core requirement for reliable behavior.

arXiv Computation and Language
5d ago

Reasoning or Rambling? Exploring the Effect of Thinking on Agent Persuasion

The paper investigates how explicit reasoning in Large Reasoning Models (LRMs) affects their ability to persuade and be persuaded. Experiments on objective and subjective tasks reveal a Persuasion Duality: reasoning boosts an agent’s persuasive power by about 21 percentage points while also making it less susceptible to incorrect persuasion by up to 10 percentage points. However, the study finds that persuasiveness often relies on superficial cues like response length and repetition rather than logical validity, and that persuasion can amplify or attenuate non‑linearly across multi‑hop agent chains. The authors also propose an attention‑guided prompt‑level adversarial argument detection method that improves agent robustness.

By Haodong Zhao, Jidong Li, Zhaomin Wu, Tianjie Ju, Zhuosheng Zhang, Bingsheng He, Gongshen Liu
arXiv Machine Learning
Jul 7

Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models

arXiv:2506. 07468v4 Announce Type: replace Abstract: Conventional large language model (LLM) safety alignment relies on a reactive, disjoint loop: attackers exploit a static model, then defenders patch exposed vulnerabilities.

By Mickel Liu, Liwei Jiang, Yancheng Liang, Simon Shaolei Du, Yejin Choi, Tim Althoff, Natasha Jaques