arXiv:2504.00285v2 Announce Type: replace
Abstract: Large Language Models (LLMs) are effective at deceiving when prompted to do so. Models that demonstrate better performance on reasoning tasks are a...
By Samuel M. Taylor, Benjamin K. Bergen
arXiv:2609.35928v1 Announce Type: cross
Abstract: Multi-agent LLM systems increasingly mix models from several providers, yet exposing each agent's underlying model identity to its peers significantl...
By Xavier Del Giudice, Alessio Palma, Matteo Migliarini, Fabio Galasso, Indro Spinelli
arXiv:2606. 27443v1 Announce Type: new Abstract: Personality prompting shapes how large language models communicate, yet whether these behavioral shifts affect objective task outcomes remains under-explored.
By Aryan Keluskar, Amrita Bhattacharjee, Huan Liu
arXiv:2608.22152v1 Announce Type: new
Abstract: Multi-agent systems built from large language models are deployed widely, yet how much performance is lost when two LLMs must coordinate rather than ac...
By Weixiang Sun, Zehong Wang, Hong Huang, Colby Nelson, Yanfang Ye
arXiv:2511. 04500v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed as decision-making agents in high-stakes domains and as imitators of human behavior in the social and behavioral sciences.
By Andrea Cera Palatsi, Samuel Martin-Gutierrez, Ana S. Cardenal, Max Pellert
arXiv:2608.30428v1 Announce Type: cross
Abstract: Strategic deception by LLM and VLM agents has emerged as a central AI alignment and safety concern. Social-deduction games (where each player holds a...
By Jaewoo Ahn, Junseo Kim, Hyunseo Kim, Heeseung Yun, Jaehyeon Son, Zsolt Kira, Gunhee Kim
arXiv:2510. 20963v2 Announce Type: replace Abstract: Multi-agent debate (MAD) was proposed as a promising approach for ensembling the wisdom of multiple large language models (LLMs) to improve reasoning and provide effective supervision to superhuman LLMs.
By Yongqiang Chen, Gang Niu, James Cheng, Bo Han, Masashi Sugiyama
arXiv:2608. 08199v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly involved in group decision-making with other LLMs and humans.
By Wenwen He, Wenke Huang, Wei Yang Bryan Lim, Dacheng Tao
The paper investigates how Large Language Model (LLM) agents can collaborate on a shared task under information asymmetry, using a table‑top version of Einstein Puzzles. It introduces a fine‑tuning‑plus‑verifier framework that equips agents with communication strategies and environmental verification signals. Results show that aligned communication is crucial for rule understanding and human trust, while a verifier improves task comprehension and promotes safer, interpretable collaboration.
By Run Peng, Ziqiao Ma, Amy Pang, Sikai Li, Zhang Xi-Jia, Yingzhuo Yu, Cristian-Paul Bara, Joyce Chai
arXiv:2604.09746v2 Announce Type: replace-cross
Abstract: As large language models (LLMs) are increasingly deployed as autonomous agents, understanding how strategic behavior emerges in multi-agent e...
By Aarush Sinha, Arion Das, Soumyadeep Nag, Charan Karnati, Shravani Nag, Chandra Vadhan Raj, Aman Chadha, Vinija Jain, Suranjana Trivedy, Amitava Das
Large language model (LLM)-based multi-agent systems (MAS) have demonstrated great potential in solving tasks with execution complexity, by distributing subtasks across cooperative agents. However, this divide-and-conquer paradigm falls short on decision-making tasks that are also prevalent in the real world.
The paper investigates why large language models (LLMs) struggle in strategic decision-making under incomplete information. It identifies two key gaps: an observation‑belief gap where LLMs’ internal representations of game states are accurate but brittle, and a belief‑action gap where converting these internal beliefs into actions is weak, leading to suboptimal payoffs. Experiments with Llama 3.1, Qwen3, and gpt‑oss confirm that acting optimally on decoded beliefs would improve outcomes in most games, highlighting a bottleneck in belief‑to‑action conversion.
By Jan Sobotka, Mustafa O. Karabag, Ufuk Topcu