Payoff scaling shapes cooperation in LLM agents across languages
arXiv:2601. 19082v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed as autonomous agents that negotiate, coordinate, and act on behalf of users.
arXiv:2607. 27536v1 Announce Type: cross Abstract: Learning a strategic task changes more than what is directly taught: fine-tuning on one game can either enhance or degrade an agent's ability to reason in another.
arXiv:2601. 19082v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed as autonomous agents that negotiate, coordinate, and act on behalf of users.
arXiv:2510. 10813v2 Announce Type: replace Abstract: Large Language Models (LLMs) are increasingly applied to domains that require reasoning about other agents' behavior, such as negotiation, policy design, and market simulation.
arXiv:2608. 07490v1 Announce Type: cross Abstract: Large language model agents are increasingly evaluated through games, but most benchmarks emphasize final outcomes rather than how players learn from repeated interaction.
Dialogue games represent a challenging setting where complex cognitive skills are required to accomplish tasks while coordinating with other players. Considering that language represents an interface...
arXiv:2609.07478v1 Announce Type: new Abstract: Large language models act as strategic agents and models of human choice, yet choosing like a strategic agent does not mean computing like one. We reco...
arXiv:2608. 12626v1 Announce Type: cross Abstract: Strategic reasoning in Large Language Models (LLMs) within long-horizon environments is often limited by inconsistent subgoals.
arXiv:2609.00474v1 Announce Type: cross Abstract: LLMs are increasingly deployed as orchestrators that coordinate specialized subagents to solve complex tasks through natural language. However, in ma...
The paper investigates whether large language models (LLMs) exhibit language‑specific skill differences by having two identical model instances play a text‑based game in different languages. Using a multilingual extension of TextArena, the authors evaluate three open‑weight models across eight languages and six games, finding that the same model can show markedly different performance—varying win–loss margins, invalid actions, and strategic choices—depending on the language interface. Analyses pinpoint language‑specific failures in spatial reasoning, card‑conditioned decisions, and optimal move selection, and demonstrate that adjusting the intermediate reasoning language can recover much of the lost performance.
Large language models access knowledge inconsistently across languages, but to what extent do they differ in their skill sets when interacting with different languages? This work quantifies cross-ling...
The paper introduces the Abstraction Agent, a zero‑shot pipeline that employs a large language model to automatically generate continuous strategic features from a natural‑language game description, score private states, and cluster them into abstraction buckets without any game‑specific evaluators or training data. The pipeline consists of four phases—feature discovery with calibration anchors, batched private‑state scoring, correlation‑based feature selection, and k‑means clustering—and achieves significant reductions in lifted‑strategy exploitability in heads‑up no‑limit Texas hold’em and outperforms scalar rank baselines in ROVER Trials. The method also transfers to other games such as four‑card Pot‑Limit Omaha, HUNL preflop and flop, and Riichi Mahjong, demonstrating that it can uncover strategic concepts that align with recognized game theory insights.
arXiv:2509. 23102v4 Announce Type: replace Abstract: Reinforcement learning from human feedback (RLHF) has emerged as the standard paradigm for aligning large language models with human preferences.
arXiv:2607. 27574v1 Announce Type: new Abstract: Activation steering has emerged in large language models as a lightweight alternative for dynamically changing a model's behavior at inference time.