arXiv:2603.18007v2 Announce Type: replace-cross
Abstract: The study explores whether current Large Language Models (LLMs) exhibit Theory of Mind (ToM) capabilities -- specifically, the ability to inf...
By Anna Babarczy, Andras Lukacs, Peter Vedres, Zeteny Bujka
CogGym is a scalable, unified framework that standardizes diverse cognitive experiments into a task‑agnostic Experiment Markup Language (EML) for systematic comparison of human and AI behavior. The initial release curates 258 experiments from 100 papers focused on human commonsense reasoning and evaluates 50 large language models, revealing a scaling trend where larger models better reproduce human judgments but still lag far behind human split‑half reliability. The framework aims to continually incorporate new cognitive science experiments to track where model behavior aligns with or diverges from human cognition as models evolve.
By Lance Ying, Jinzhou Wu, Yingshan Susan Wang, Shivam Aarya, Luca M. Schulze Buschoff, Harry Chen, Katherine M. Collins, Andrea de Varda, Shuhao Fu, Sean Dae Houlihan, Akshay K. Jagadish, Guangyuan Jiang, Samuel Kiegeland, Tetsu Kurumisawa, Rongzhi Liu, Ryan Liu, Ningshan Ma, Kathryn McGregor, Younes Strittmatter, Polina Tsvilodub, Jacob Hoover Vigly, Sarah Wu, Enjie Xu, Yiling Yun, Kelsey Allen, Tyler Brooke-Wilson, Brian Christian, Evelina Fedorenko, Michael C. Frank, Michael Franke, Tao Gao, Samuel J. Gershman, Robert D. Hawkins, Jennifer Hu, Julian Jara-Ettinger, Max Kleiman-Weiner, Sydney Levine, Tal Linzen, Hongjing Lu, Timothy O'Donnell, Desmond C. Ong, Steven T. Piantadosi, Rebecca Saxe, Eric Schulz, Tianmin Shu, Felix A. Sosa, Ilia Sucholutsky, Tan Zhi-Xuan, Tomer Ullman, Fei Xu, Ilker Yildirim, Jian-Qiao Zhu, Thomas L. Griffiths, Tobias Gerstenberg, Kevin Smith, Joshua B. Tenenbaum
arXiv:2609.22152v1 Announce Type: new
Abstract: Imagination performs as a high-level function of large language models (LLMs) which determines the potential of how an LLM creates unseen or creative c...
By Zixuan Tang, Hongzong Li, Shuxin Zhuang, Dapeng Wu, Zi Liang
arXiv:2606. 30481v1 Announce Type: cross Abstract: Current large language models are extraordinary statistical engines.
By Ziqin Yuan, Jaymari Chua
arXiv:2502. 20502v2 Announce Type: replace Abstract: Recent advances in Artificial Intelligence (AI) have yielded powerful computational models that, by learning from vast amounts of human-generated data, are increasingly posited as approximate models of human cognition.
By Lance Ying, Katherine M. Collins, Lionel Wong, Ilia Sucholutsky, Ryan Liu, Adrian Weller, Tianmin Shu, Thomas L. Griffiths, Joshua B. Tenenbaum
arXiv:2609.08003v1 Announce Type: new
Abstract: Behavioral foundation models have been proposed as stand-ins for human participants across settings, but it is unclear whether theories discovered on t...
By Akshay K. Jagadish, Younes Strittmatter, Nori Jacoby, Eric Schulz, Nathaniel Daw, Thomas L. Griffiths, Suyog H. Chandramouli
arXiv:2607. 15883v1 Announce Type: cross Abstract: Large language models are broadly capable, yet in sustained one-to-one conversation they still read as flat: competent, responsive, and somehow not quite the presence of a mind.
By Sebastian Cochinescu
arXiv:2605.06524v3 Announce Type: replace
Abstract: Reliable human-machine discrimination is becoming increasingly important as Large Language Models and autonomous agents are deployed in online sett...
By Milena Rmus, Mathew D. Hardy, Thomas L. Griffiths, Mayank Agrawal
The paper investigates why large language models (LLMs) produce hallucinations—outputs that are fabricated, unverifiable, or contradictory to source material—and argues that these hallucinations have philosophical implications for machine consciousness. It reviews known causes such as source‑target divergence, training‑inference discrepancies, and overfitting, and presents two empirical studies: one showing that higher temperature settings in GPT models yield plausible but incorrect answers, while lower temperatures produce accurate ones; and another demonstrating that an encoder‑only model trained on encyclopedic data answers factually without embellishment, suggesting hallucinations arise from exposure to subjective, socially diverse data rather than cognitive ability. Drawing on Turing, Searle’s Chinese Room, the frame problem, and cybernetic theory, the authors contend that a model’s self‑reports of emotion or sentience fall within the definition of hallucination, implying that any future machine consciousness may remain epistemically inaccessible because it would be indistinguishable from an advanced hallucination.
By Kristina \v{S}ekrst
arXiv:2510. 02660v2 Announce Type: replace-cross Abstract: When researchers claim AI systems possess ToM or mental models, they are fundamentally discussing behavioral predictions and bias corrections rather than genuine mental states.
By Xiaoyun Yin, Elmira Zahmat Doost, Shiwen Zhou, Garima Arya Yadav, Jamie C. Gorman
arXiv:2509. 14474v3 Announce Type: replace Abstract: The debate around Artificial General Intelligence (AGI) remains open due to two fundamentally different goals: replicating human-level performance versus replicating human-like cognitive processes.
By Meltem Subasioglu, Nevzat Subasioglu
The paper defends the 'Whole Hog Thesis', arguing that sophisticated large language models such as ChatGPT are full linguistic and cognitive agents, possessing understanding, beliefs, desires, knowledge, and intentions. It rejects low‑level computational starting points and instead builds its case from high‑level behavioral observations, using Holistic Network Assumptions to link actions to mental states. The authors systematically rebut common objections—such as hallucinations and planning errors—by showing these resemble human fallibility and by challenging the necessity of traditional conditions like embodiment or semantic grounding.
By Herman Cappelen, Josh Dever