arXiv AI

Can machines think efficiently?

The article proposes an updated Turing Test that incorporates energy consumption as a key metric, arguing that the original test is insufficient for distinguishing human from machine intelligence in the context of modern AI. It suggests that by adding an energy constraint, the test evaluates intelligence through the lens of efficiency, linking abstract thinking to tangible resource limits. The new test also provides a measurable, practical endpoint, encouraging society to balance AI time savings against total resource costs.

arXiv AI
Sep 10

Turing's First Imitation Game: Design Concepts and a Human-Approximates-Machine Reading

The paper analyzes Turing’s 1948 report "Intelligent Machinery" as a foundational source for later imitation games, highlighting key design concepts such as the possibility of machine errors, the exclusion of irrelevant physical traits, the role of a human judge, and Turing’s view that intellectual activity is largely search. It argues that limiting the human contestant to a weak chess player heightens the importance of intellectual search, making human behavior more comparable to machine behavior. This reframes the 1948 game as a human‑approximates‑machine scenario, suggesting that imitation games can probe when human intelligence becomes machine‑like under specific task constraints.

By Sharon Temtsin, Christoph Bartneck
arXiv AI
Aug 13

On Benchmarking Human-Like Intelligence in Machines

arXiv:2502. 20502v2 Announce Type: replace Abstract: Recent advances in Artificial Intelligence (AI) have yielded powerful computational models that, by learning from vast amounts of human-generated data, are increasingly posited as approximate models of human cognition.

By Lance Ying, Katherine M. Collins, Lionel Wong, Ilia Sucholutsky, Ryan Liu, Adrian Weller, Tianmin Shu, Thomas L. Griffiths, Joshua B. Tenenbaum
arXiv AI
Sep 21

CogGym: Towards Large-Scale Comparative Evaluation of Human and Machine Cognition

CogGym is a scalable, unified framework that standardizes diverse cognitive experiments into a task‑agnostic Experiment Markup Language (EML) for systematic comparison of human and AI behavior. The initial release curates 258 experiments from 100 papers focused on human commonsense reasoning and evaluates 50 large language models, revealing a scaling trend where larger models better reproduce human judgments but still lag far behind human split‑half reliability. The framework aims to continually incorporate new cognitive science experiments to track where model behavior aligns with or diverges from human cognition as models evolve.

By Lance Ying, Jinzhou Wu, Yingshan Susan Wang, Shivam Aarya, Luca M. Schulze Buschoff, Harry Chen, Katherine M. Collins, Andrea de Varda, Shuhao Fu, Sean Dae Houlihan, Akshay K. Jagadish, Guangyuan Jiang, Samuel Kiegeland, Tetsu Kurumisawa, Rongzhi Liu, Ryan Liu, Ningshan Ma, Kathryn McGregor, Younes Strittmatter, Polina Tsvilodub, Jacob Hoover Vigly, Sarah Wu, Enjie Xu, Yiling Yun, Kelsey Allen, Tyler Brooke-Wilson, Brian Christian, Evelina Fedorenko, Michael C. Frank, Michael Franke, Tao Gao, Samuel J. Gershman, Robert D. Hawkins, Jennifer Hu, Julian Jara-Ettinger, Max Kleiman-Weiner, Sydney Levine, Tal Linzen, Hongjing Lu, Timothy O'Donnell, Desmond C. Ong, Steven T. Piantadosi, Rebecca Saxe, Eric Schulz, Tianmin Shu, Felix A. Sosa, Ilia Sucholutsky, Tan Zhi-Xuan, Tomer Ullman, Fei Xu, Ilker Yildirim, Jian-Qiao Zhu, Thomas L. Griffiths, Tobias Gerstenberg, Kevin Smith, Joshua B. Tenenbaum