arXiv:2604. 11840v3 Announce Type: replace-cross Abstract: Language models are increasingly used to simulate people: survey respondents, negotiators, stakeholders in policy exercises.
By Sandro Andric
The paper investigates why large language models (LLMs) struggle in strategic decision-making under incomplete information. It identifies two key gaps: an observation‑belief gap where LLMs’ internal representations of game states are accurate but brittle, and a belief‑action gap where converting these internal beliefs into actions is weak, leading to suboptimal payoffs. Experiments with Llama 3.1, Qwen3, and gpt‑oss confirm that acting optimally on decoded beliefs would improve outcomes in most games, highlighting a bottleneck in belief‑to‑action conversion.
By Jan Sobotka, Mustafa O. Karabag, Ufuk Topcu
arXiv:2609.25686v1 Announce Type: cross
Abstract: Long-horizon assigned work requires an LLM agent to track the state of a task: which steps are done, blocked, cancelled, or open to repetition. Agent...
By Chenyu Zhang, Wonbin Kweon, Jiawei Han
The paper introduces XiangqiBench, an executable benchmark for evaluating large language model agents in Chinese chess. It tracks 8,568 multi‑turn trajectories from 12 frontier LLMs, revealing that metrics such as the Conversion Gap, Consistency Gap, and Simulation Gap overstate true closed‑loop success. The study shows that simply naming the correct move is insufficient; agents must reliably carry a plan through to a verified outcome against an opponent.
By Yekun Chai, Qiwei Peng, Haoyi Xiong
arXiv:2606. 19494v1 Announce Type: new Abstract: Multi-agent LLM deliberation, where agents exchange and revise answers over several rounds, is increasingly used to improve reasoning and accuracy, yet how and why it works is rarely modelled.
By Apurba Pokharel, Ram Dantu
arXiv:2606. 16014v1 Announce Type: cross Abstract: Many games rely on storytelling combined with systems that track levelling, NPC behaviour, and consequence simulation; bridging tightly-authored narrative with deeply-simulated worlds -- most acute in sandbox and open-world settings -- has been prohibitively expensive.
By Yuhang Huang, Chenmiao Li, Chaowei Fang
arXiv:2608. 04289v1 Announce Type: new Abstract: Long-horizon agents increasingly use persistent memory and tools to take actions with external side effects.
By Mayur Akewar, Ravi Ranjan
The study shows that large language model (LLM) agents are far more likely to commit to a directional prediction when presented with a professional‑looking market panel than when asked the same question directly, with commitment rates rising from 6.5% to 54.0% across 12 frontier models. Even when the panel’s data is entirely fabricated, commitment still increases significantly, indicating that the authority of the presentation, rather than the truth of the information, drives confident action. The authors demonstrate that this act/don’t‑act decision gate is narrow, model‑specific, and can be mitigated through supervised fine‑tuning, though its effectiveness depends on response format and context.
whyItMatters":"The findings reveal a specific vulnerability in LLMs where presentation style can override factual accuracy, highlighting the need for careful design and training to prevent misleading confidence in uncertain scenarios."
By Pranav Aggarwal
arXiv:2510. 10813v2 Announce Type: replace Abstract: Large Language Models (LLMs) are increasingly applied to domains that require reasoning about other agents' behavior, such as negotiation, policy design, and market simulation.
By Enric Junque de Fortuny, Veronica Roberta Cappelli
Agentic systems increasingly gate actions on a model's own stated confidence, which assumes confidence tracks correctness at the moment of acting. We test this in a hidden-information chess variant wh...
arXiv:2509. 17192v3 Announce Type: replace Abstract: LLM-based social simulations can make a generated transcript look like a single behavioral signal, but the model behind that transcript may be doing several different jobs: choosing what an actor says or does, deciding what happens after an action, or both.
By Glenn Matlin, Isaac Song, Yixiong Hao, Parv Mahajan, Evan Montoya, Ryan Bard, Stuart R. Topp, Anthony Wen-Ming Zang, Mohammed Rehan Parwani, Soham Shetty, Mark Riedl
The paper investigates whether large language models (LLMs) correctly gauge their confidence when acting in a hidden‑information chess variant. In experiments where the location of a hidden royal piece is repeatedly relocated, the models’ stated probabilities about the piece’s position were almost never accurate at high confidence levels, with a calibration deficit concentrated in those high‑confidence events. Across multiple model configurations and providers, the same pattern emerged, and conventional evaluation metrics such as legality, cost, latency, and completion rate were found to be uncorrelated with belief quality, yet a model could still win the game despite poor confidence estimates.
By Bhushan Kashinath Joshi