KC-Bench is a dynamic interactive benchmark designed to evaluate how large language model agents reconcile user instructions, internal knowledge, and real‑time environmental observations. It contains 238 manually curated multi‑turn tasks that test world‑knowledge conflicts, input inconsistencies, and multi‑source temporal conflicts, using a user simulator, stateful tools, deterministic environment assertions, an open‑source natural‑language evaluator, and human trajectory verification. Evaluation of nine models—including DeepSeek‑V4‑Flash, GLM‑5.2, and MiniMax‑M3—reveals significant cross‑domain variation, with no model reliably handling factual correction, identity consistency checking, and temporal conflict resolution across all settings, and shows that missed conflicts can propagate to tool calls or synthetic protected‑data flows.
By Yaxing Lyu, Shengjie Zhou, Binbin Toh, Pengyu Zhu, Lijun Li
arXiv:2607. 06757v1 Announce Type: new Abstract: Agent-based modeling (ABM) has the capability to model millions of individuals and their interactions, which is useful for policy making.
By Sifat Afroj Moon, Dakotah Maguire, Adam Spannaus, Joe Tuccillo, Maksudul Alam, Sudip K. Seal, John Gounley, Heidi Hanson
arXiv:2606. 14715v1 Announce Type: cross Abstract: LLM agents are increasingly used to simulate real world interactions, but it remains unclear whether simulated behaviors preserve the content patterns and interaction dynamics of real human behaviors.
By Yaoning Yu, Ye Yu, Haojing Luo, Haohan Wang
arXiv:2606. 08200v1 Announce Type: new Abstract: Evaluating LLM-powered interactive social agents is challenging because socially relevant behaviors depend not only on isolated outputs, but also on prior interactions, social roles, and downstream actions.
By Hyogon Ryu, Jeonghwan Kim, Yewon Lim, Chaeun Lee, Jeongwook Kim, Donghoon Ham
KC-Bench is a dynamic interactive benchmark designed to evaluate how large language model agents reconcile user instructions, internal knowledge, and real‑time environmental observations. It consists of 238 manually curated multi‑turn tasks that test world‑knowledge conflicts, input inconsistencies, and multi‑source temporal conflicts, using a user simulator, stateful tools, deterministic environment assertions, an open‑source natural‑language evaluator, and human trajectory verification. Evaluation of nine models—including DeepSeek‑V4‑Flash, GLM‑5.2, and MiniMax‑M3—reveals significant cross‑domain variation, with no model reliably handling factual correction, identity consistency, and temporal conflict resolution across all settings, and shows that missed conflicts can propagate to tool calls or synthetic protected‑data flows.
arXiv:2607. 00989v1 Announce Type: cross Abstract: Semantic trajectory analysis has recently emerged as an approach for modeling human movement by capturing implicit patterns and behaviors through semantic information (e.
By Ziyue Lin, Xinhang Xie, Kangyi Wang, Siming Chen