arXiv:2509. 00559v3 Announce Type: replace Abstract: Humans intuitively navigate social interactions by simulating unspoken dynamics and reasoning about others' perspectives, even with limited information.
By Xuhui Zhou, Jiarui Liu, Akhila Yerukola, Hyunwoo Kim, Maarten Sap
arXiv:2512. 09706v2 Announce Type: replace Abstract: The paradigm of agentic AI is shifting from engineered complex workflows to post-training native models.
By Kaichen He, Zihao Wang, Muyao Li, Anji Liu, Yitao Liang
arXiv:2607. 14485v1 Announce Type: new Abstract: Large language model (LLM)-based generative agents simulate human behavior through long-horizon decision-making processes that comprise intermediate steps such as planning, memory retrieval, reflection, and action selection.
By Wenchang Gao, Pingyue Sheng, Lanlan Qiu, Yunfei Ma, Jian Zhao, Baicheng Chen, Kangda Wang, Yuyang Tian, Shunqiang Mao, Tianxing He
arXiv:2606. 18537v1 Announce Type: new Abstract: Humans often acquire new skills by observing others, since observed behaviors implicitly reveal how to act in an environment.
By Caleb Chang, Davin Win Kyi, Natasha Jaques, Karen Leung
The paper presents UBCL, a reinforcement learning framework that generates controllable and diverse player behaviors without using human gameplay data. By defining behavior in an N‑dimensional continuous space and training a single PPO‑based multi‑agent policy with target behavior vectors, the method learns how actions affect behavioral statistics such as aggressiveness, mobility, and cooperativeness. Experiments in a custom Unity multiplayer game demonstrate that UBCL achieves greater behavioral diversity than a win‑only baseline and accurately matches specified behavior vectors across a range of targets.
By Atahan Cilan, Atay \"Ozg\"ovde
UnifiedPlayers is a cooperative framework that jointly adapts planning, execution, and evaluation for tool-integrated reinforcement learning agents. It consists of a Planning Player that generates tasks, an Execution Player that creates multi-turn trajectories with Python tool calls, and an Evaluation Player that builds executable verifiers, all coordinated by role‑specific rewards under GRPO. The approach outperforms prior baselines on mathematical and general reasoning benchmarks and yields a verifier with high adversarial detection accuracy and more discriminative reward signals.
By Wenjie Liao, Liangjie Zhao, Zehong Cao