arXiv:2501. 17629v2 Announce Type: replace-cross Abstract: Several studies claim that large language models have passed the Turing Test and hence can "think", yet none follow Turing's original instructions precisely.
By Sharon Temtsin, Diane Proudfoot, David Kaber, Christoph Bartneck
arXiv:2608. 05558v1 Announce Type: cross Abstract: This paper examines Turing's 1948 report, "Intelligent Machinery", as an important conceptual source for the later imitation games.
By Sharon Temtsin, Christoph Bartneck
The paper analyzes Turing’s 1948 report "Intelligent Machinery" as a foundational source for later imitation games, highlighting key design concepts such as the possibility of machine errors, the exclusion of irrelevant physical traits, the role of a human judge, and Turing’s view that intellectual activity is largely search. It argues that limiting the human contestant to a weak chess player heightens the importance of intellectual search, making human behavior more comparable to machine behavior. This reframes the 1948 game as a human‑approximates‑machine scenario, suggesting that imitation games can probe when human intelligence becomes machine‑like under specific task constraints.
By Sharon Temtsin, Christoph Bartneck
arXiv:2502. 20502v2 Announce Type: replace Abstract: Recent advances in Artificial Intelligence (AI) have yielded powerful computational models that, by learning from vast amounts of human-generated data, are increasingly posited as approximate models of human cognition.
By Lance Ying, Katherine M. Collins, Lionel Wong, Ilia Sucholutsky, Ryan Liu, Adrian Weller, Tianmin Shu, Thomas L. Griffiths, Joshua B. Tenenbaum
CogGym is a scalable, unified framework that standardizes diverse cognitive experiments into a task‑agnostic Experiment Markup Language (EML) for systematic comparison of human and AI behavior. The initial release curates 258 experiments from 100 papers focused on human commonsense reasoning and evaluates 50 large language models, revealing a scaling trend where larger models better reproduce human judgments but still lag far behind human split‑half reliability. The framework aims to continually incorporate new cognitive science experiments to track where model behavior aligns with or diverges from human cognition as models evolve.
By Lance Ying, Jinzhou Wu, Yingshan Susan Wang, Shivam Aarya, Luca M. Schulze Buschoff, Harry Chen, Katherine M. Collins, Andrea de Varda, Shuhao Fu, Sean Dae Houlihan, Akshay K. Jagadish, Guangyuan Jiang, Samuel Kiegeland, Tetsu Kurumisawa, Rongzhi Liu, Ryan Liu, Ningshan Ma, Kathryn McGregor, Younes Strittmatter, Polina Tsvilodub, Jacob Hoover Vigly, Sarah Wu, Enjie Xu, Yiling Yun, Kelsey Allen, Tyler Brooke-Wilson, Brian Christian, Evelina Fedorenko, Michael C. Frank, Michael Franke, Tao Gao, Samuel J. Gershman, Robert D. Hawkins, Jennifer Hu, Julian Jara-Ettinger, Max Kleiman-Weiner, Sydney Levine, Tal Linzen, Hongjing Lu, Timothy O'Donnell, Desmond C. Ong, Steven T. Piantadosi, Rebecca Saxe, Eric Schulz, Tianmin Shu, Felix A. Sosa, Ilia Sucholutsky, Tan Zhi-Xuan, Tomer Ullman, Fei Xu, Ilker Yildirim, Jian-Qiao Zhu, Thomas L. Griffiths, Tobias Gerstenberg, Kevin Smith, Joshua B. Tenenbaum
arXiv:2608. 16213v1 Announce Type: new Abstract: Intelligence is constituted by \textit{process} (iterative activity through which output emerges), not in the output itself.
By Michael J. Richardson, Ayeh Alhasan, Cassandra Crone, M. Paula Diaz Monfort, Patrick Nalepka, Mark Dras, Rachel W. Kallen, David M. Kaplan
arXiv:2605.10851v2 Announce Type: replace
Abstract: We initiate the study of the Generalized Turing Test (GTT), a formal generalization of Turing's imitation game from humans to arbitrary interactive...
By Daniel Mitropolsky, Riccardo Neumarker, Emanuele Rimoldi, Susan S. Hong, Tomaso Poggio
arXiv:2510. 02660v2 Announce Type: replace-cross Abstract: When researchers claim AI systems possess ToM or mental models, they are fundamentally discussing behavioral predictions and bias corrections rather than genuine mental states.
By Xiaoyun Yin, Elmira Zahmat Doost, Shiwen Zhou, Garima Arya Yadav, Jamie C. Gorman
arXiv:2605.06524v3 Announce Type: replace
Abstract: Reliable human-machine discrimination is becoming increasingly important as Large Language Models and autonomous agents are deployed in online sett...
By Milena Rmus, Mathew D. Hardy, Thomas L. Griffiths, Mayank Agrawal
arXiv:2608. 14407v1 Announce Type: new Abstract: We present a survey of the past and future of AI Scientists: machines capable of automating science.
By Ross D. King
arXiv:2606. 30481v1 Announce Type: cross Abstract: Current large language models are extraordinary statistical engines.
By Ziqin Yuan, Jaymari Chua
arXiv:2606. 12848v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for tasks once reserved for trained researchers, including hypothesis generation, specification choice, and drafting conclusions.
By Chen Zhu, Xiaolu Wang, Weilong Zhang