arXiv:2609.38604v1 Announce Type: cross
Abstract: Modern LLM agents increasingly tackle complex tasks through interactive, long-horizon exchanges with users, while existing benchmarks generally assum...
By Zheyuan Zhang, Mengyuan Chao, Ke Xiao, Ziyi Chen, Daoan Zhang, Yan Zhang, Yanfang Ye, Wei Xu
arXiv:2606. 05304v1 Announce Type: new Abstract: Multi-agent systems (MAS) built on large language models are typically organized around roles, pipelines, and turn schedules, while the content that agents pass to one another is often left as unconstrained natural language.
By Chen Huang, Yuhao Wu, Wenxuan Zhang
arXiv:2608.22055v1 Announce Type: new
Abstract: Suppose one embodied agent knows what must be built, while its teammate alone knows which transformation its workcell can perform. Neither local view d...
By Peng He, Junning Zhu, Haohan Yuan, Jianpeng Liang
The study examines how Theory of Mind (ToM) reasoning and prosocial beliefs influence large language models (LLMs) in the ultimatum game. By initializing LLM agents with Greedy, Fair, or Selfless beliefs and applying chain‑of‑thought or varying levels of ToM reasoning, the authors ran 2,700 simulations across several models, including o3‑mini and DeepSeek‑R1 Distilled Qwen 32B. Results show that ToM‑enhanced LLMs align more closely with human decision patterns, exhibit greater consistency, and achieve better negotiation outcomes, with Llama 3.3 70B producing the most belief‑consistent reasoning.
whyItMatters":"The findings clarify the importance of incorporating Theory of Mind into LLMs to improve their alignment with human norms in cooperative decision‑making tasks."
By Neemesh Yadav, Yihuai Lan, Shan Dong, Mai Hieu Hien, Palakorn Achananuparp, Jing Jiang, Ee-Peng Lim
arXiv:2608. 02578v1 Announce Type: cross Abstract: World Action Models (WAMs) augment robot policies with action-conditioned predicted futures, but a plausible future alone does not justify changing the action that a bimanual policy would execute.
By Shuaijun Liu, Qifu Wen, Shuyang Hao, Qi Luo, Chenglong Zhang, Feiyang You, Chengyu Wu, Ningxin Su
arXiv:2609.22245v1 Announce Type: new
Abstract: Large language models can produce fluent explanations for chess moves, but plausible language does not necessarily reflect the reasoning behind a decis...
By Angelina Parfenova
arXiv:2607. 20734v1 Announce Type: new Abstract: As LLMs become more capable, they are increasingly deployed as collaborative agents, taking on user-delegated tasks through iterative interaction.
By Jihoon Tack, Philippe Laban, Jennifer Neville
The paper introduces a framework for evaluating how large language model agents revise their success criteria after failures, defining five non‑compensatory conditions that must be met for a criterion revision to be considered valid. Using the CMB‑0.1 protocol, the authors test twelve cross‑domain scenarios across four system configurations, finding that no model trial satisfies all five conditions and highlighting specific failure modes such as zero‑state reconstruction and inadequate intervention sensitivity. They propose a more stringent trace‑anchored CMB‑0.4 protocol to better isolate and measure criterion revision in future studies.
By Guodong Xu
The paper critiques the standard fixed-rule approach for deriving labels from human feedback in human-robot collaboration, showing that human-provided implication labels often differ and improve reward learning. It introduces IMPLIED, a method that starts with fixed-rule implications but learns to infer and revise accepted/rejected action labels over time, outperforming both the fixed rule and LLM baselines on recorded trajectories and a physical pizza‑making study. As a result, IMPLIED reduces preference‑estimation error and yields robot actions that better align with combined reward objectives.
By Qiping Zhang, Kate Candon, Debasmita Ghose, Marynel V\'azquez
arXiv:2607. 03598v1 Announce Type: cross Abstract: When a person shares something with a language model, the model often answers the surface of the message rather than what the sender was doing by sending it: share a finished project and it critiques the code; share a raw late-night line and it runs a wellness check.
By Alex Kwon
Translating natural-language planning intent into verified plans is a longstanding challenge: people communicate goals in language, while classical planners require formal PDDL specifications. Recent agentic frameworks bridge this gap by orchestrating a pool of specialized repair agents inside a verifier-checked refinement loop, but the orchestrator at the centre is itself a prompted frontier LLM, paying a frontier-LLM API call at every refinement step.
The paper introduces CAPA, a Collaborative Agent Predictive Architecture designed to give large language model (LLM) agents situational awareness in online meetings. CAPA uses a Perceiver to update meeting state, a Predictor to forecast conversation flow, a Controller to decide speaking actions, and a Generator to phrase contributions. Evaluated on 137 AMI meetings, CAPA reduces the silence rate from 51.4% to 2.5%, doubles credited recovery, and maintains low hallucination, demonstrating that structured state tracking is key to effective delegation.
By Muneeb Khan, Frederic Kirstein, Terry Ruas, Bela Gipp