arXiv:2602.00685v2 Announce Type: replace
Abstract: Large language models (LLMs) are increasingly used as simulated participants in social science experiments, but their behavior is often unstable an...
By Xuan Liu, Haoyang Shang, Zizhang Liu, Xinyan Liu, Yunze Xiao, Yiwen Tu, Haojian Jin
arXiv:2606. 11217v1 Announce Type: cross Abstract: The proliferation of large language models (LLMs) and autonomous AI agents has given rise to a rapidly growing methodological paradigm: "in silico" behavioral experiments.
By Michelle Vaccaro
arXiv:2608. 07494v1 Announce Type: cross Abstract: AI tools like ChatGPT and DeepSeek, powered by Large Language Models (LLMs), allow users to obtain instant and effective content responses simply by typing requests, such as ``plan a three-day Vienna trip'', ``solve the attached mathematical problem'', ``draft an email to inquire review progress'', etc.
By Yiqun Zhang, Yunfan Zhang, Mingjie Zhao, Sen Feng, Yiu-ming Cheung
arXiv:2509.08494v2 Announce Type: replace-cross
Abstract: As humans delegate more tasks and decisions to artificial intelligence (AI), we risk losing control of our individual and collective futures....
By Benjamin Sturgeon, Daniel Samuelson, Jacob Haimes, Jacy Reese Anthis
Pairit is an online platform that enables researchers to design, test, and deploy live experiments on human-AI collaboration. Using a single YAML configuration file, users can specify an executable experiment graph—including pages, routing, randomization, matchmaking, chat, shared workspaces, server-hosted agents, surveys, timers, and custom HTML components—and combine any number of humans and AI agents in real-time sessions. The platform has been validated through multiple live deployments, including peer-reviewed studies, and captures high-resolution process traces of communication, negotiation, and collaborative work in human-AI dyads.
By Harang Ju, Sinan Aral
arXiv:2607. 08285v1 Announce Type: new Abstract: Current AI evaluation frameworks focus primarily on technical performance, including accuracy, robustness, reasoning ability, and policy compliance.
By Marcos Economides, Paul M. Sacher, Samuel Salzer, Alexis Michelle Abellar, Fendi Tsim, Antoine Ferr\`ere