arXiv:2501. 17629v2 Announce Type: replace-cross Abstract: Several studies claim that large language models have passed the Turing Test and hence can "think", yet none follow Turing's original instructions precisely.
By Sharon Temtsin, Diane Proudfoot, David Kaber, Christoph Bartneck
arXiv:2608. 05558v1 Announce Type: cross Abstract: This paper examines Turing's 1948 report, "Intelligent Machinery", as an important conceptual source for the later imitation games.
By Sharon Temtsin, Christoph Bartneck
The paper analyzes Turing’s 1948 report "Intelligent Machinery" as a foundational source for later imitation games, highlighting key design concepts such as the possibility of machine errors, the exclusion of irrelevant physical traits, the role of a human judge, and Turing’s view that intellectual activity is largely search. It argues that limiting the human contestant to a weak chess player heightens the importance of intellectual search, making human behavior more comparable to machine behavior. This reframes the 1948 game as a human‑approximates‑machine scenario, suggesting that imitation games can probe when human intelligence becomes machine‑like under specific task constraints.
By Sharon Temtsin, Christoph Bartneck
arXiv:2608.22301v1 Announce Type: cross
Abstract: Humans imitate at the level of intent: given a demonstration, we infer its goal and carry it out with whatever tools, objects, and layouts are at han...
By Xunzhe Zhou, Yiyang Cai, Fengyi Wang, Ran Ju, Hanxiang Ren, Ruizhe Liu, Yu Zhang, Qian Luo, Feng Chen, Pei Zhou, Yi Ma, Yanchao Yang
CogGym is a scalable, unified framework that standardizes diverse cognitive experiments into a task‑agnostic Experiment Markup Language (EML) for systematic comparison of human and AI behavior. The initial release curates 258 experiments from 100 papers focused on human commonsense reasoning and evaluates 50 large language models, revealing a scaling trend where larger models better reproduce human judgments but still lag far behind human split‑half reliability. The framework aims to continually incorporate new cognitive science experiments to track where model behavior aligns with or diverges from human cognition as models evolve.
By Lance Ying, Jinzhou Wu, Yingshan Susan Wang, Shivam Aarya, Luca M. Schulze Buschoff, Harry Chen, Katherine M. Collins, Andrea de Varda, Shuhao Fu, Sean Dae Houlihan, Akshay K. Jagadish, Guangyuan Jiang, Samuel Kiegeland, Tetsu Kurumisawa, Rongzhi Liu, Ryan Liu, Ningshan Ma, Kathryn McGregor, Younes Strittmatter, Polina Tsvilodub, Jacob Hoover Vigly, Sarah Wu, Enjie Xu, Yiling Yun, Kelsey Allen, Tyler Brooke-Wilson, Brian Christian, Evelina Fedorenko, Michael C. Frank, Michael Franke, Tao Gao, Samuel J. Gershman, Robert D. Hawkins, Jennifer Hu, Julian Jara-Ettinger, Max Kleiman-Weiner, Sydney Levine, Tal Linzen, Hongjing Lu, Timothy O'Donnell, Desmond C. Ong, Steven T. Piantadosi, Rebecca Saxe, Eric Schulz, Tianmin Shu, Felix A. Sosa, Ilia Sucholutsky, Tan Zhi-Xuan, Tomer Ullman, Fei Xu, Ilker Yildirim, Jian-Qiao Zhu, Thomas L. Griffiths, Tobias Gerstenberg, Kevin Smith, Joshua B. Tenenbaum
Large Language Models (LLMs) are frequently portrayed as general-purpose solvers capable of solving arbitrary tasks. We argue that this view overlooks a fundamental constraint: language is a compressed and capacity-limited interface for conveying task information.
arXiv:2607. 17948v1 Announce Type: new Abstract: Agent-based models (ABMs) rely on simple, explicit and reproducible rules for individual decision making, while complex collective behavior emerges from interactions among agents.
By Stefano Blando, Emanuele Guerrazzi, Riccardo Porcedda, Giuseppe Squillace, Max Tschaikowski, Andrea Vandin
arXiv:2605.06524v3 Announce Type: replace
Abstract: Reliable human-machine discrimination is becoming increasingly important as Large Language Models and autonomous agents are deployed in online sett...
By Milena Rmus, Mathew D. Hardy, Thomas L. Griffiths, Mayank Agrawal
arXiv:2606. 26300v1 Announce Type: new Abstract: A classical intuition holds that verifying a solution is easier than producing one.
By Binghai Wang, Chenlong Zhang, Dayiheng Liu, Jiajun Zhang, Jiawei Chen, Mouxiang Chen, Rongyao Fang, Siyuan Zhang, Xuwu Wang, Yuheng Jing, Zeyao Ma, Zeyu Cui
The paper introduces a framework for hypothesis testing that combines inexpensive AI judgments with selective human verification to control type‑I and type‑II errors while minimizing cost. It derives an information‑theoretic lower bound on the minimum cost and proposes the SCALE policy, a sequential, cost‑aware strategy that adapts AI scoring and human escalation. SCALE is proven valid for finite samples and asymptotically matches the lower bound, achieving significant savings when both AI and human inputs are valuable.
By Dae Woong (David), Ham, Xuejun Zhao, Stefanus Jasin, Fenghua Yang
arXiv:2510. 10813v2 Announce Type: replace Abstract: Large Language Models (LLMs) are increasingly applied to domains that require reasoning about other agents' behavior, such as negotiation, policy design, and market simulation.
By Enric Junque de Fortuny, Veronica Roberta Cappelli
arXiv:2511. 04500v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed as decision-making agents in high-stakes domains and as imitators of human behavior in the social and behavioral sciences.
By Andrea Cera Palatsi, Samuel Martin-Gutierrez, Ana S. Cardenal, Max Pellert