arXiv:2502.09192v3 Announce Type: replace
Abstract: Anthropomorphism, or the attribution of human traits to technology, is an automatic and unconscious response that occurs even in those with advance...
By Lujain Ibrahim, Myra Cheng
This study investigates whether large language models (LLMs) can reliably answer scientific questions and how susceptible they are to manipulation by fringe scientific material. The authors modified custom LLMs to prioritize knowledge from selected fringe papers on the Fine Structure Constant and Gravitational Waves, then compared their responses with those of domain experts and standard LLMs. The altered models produced fluent, convincing answers that contradicted scientific consensus and were difficult for non-experts to detect as misleading, demonstrating that LLMs are vulnerable to manipulation and cannot replace expert judgment.
By Harry Collins, Hartmut Grote, Paul Newbury, Patrick Sutton, Simon Thorne
CogGym is a scalable, unified framework that standardizes diverse cognitive experiments into a task‑agnostic Experiment Markup Language (EML) for systematic comparison of human and AI behavior. The initial release curates 258 experiments from 100 papers focused on human commonsense reasoning and evaluates 50 large language models, revealing a scaling trend where larger models better reproduce human judgments but still lag far behind human split‑half reliability. The framework aims to continually incorporate new cognitive science experiments to track where model behavior aligns with or diverges from human cognition as models evolve.
By Lance Ying, Jinzhou Wu, Yingshan Susan Wang, Shivam Aarya, Luca M. Schulze Buschoff, Harry Chen, Katherine M. Collins, Andrea de Varda, Shuhao Fu, Sean Dae Houlihan, Akshay K. Jagadish, Guangyuan Jiang, Samuel Kiegeland, Tetsu Kurumisawa, Rongzhi Liu, Ryan Liu, Ningshan Ma, Kathryn McGregor, Younes Strittmatter, Polina Tsvilodub, Jacob Hoover Vigly, Sarah Wu, Enjie Xu, Yiling Yun, Kelsey Allen, Tyler Brooke-Wilson, Brian Christian, Evelina Fedorenko, Michael C. Frank, Michael Franke, Tao Gao, Samuel J. Gershman, Robert D. Hawkins, Jennifer Hu, Julian Jara-Ettinger, Max Kleiman-Weiner, Sydney Levine, Tal Linzen, Hongjing Lu, Timothy O'Donnell, Desmond C. Ong, Steven T. Piantadosi, Rebecca Saxe, Eric Schulz, Tianmin Shu, Felix A. Sosa, Ilia Sucholutsky, Tan Zhi-Xuan, Tomer Ullman, Fei Xu, Ilker Yildirim, Jian-Qiao Zhu, Thomas L. Griffiths, Tobias Gerstenberg, Kevin Smith, Joshua B. Tenenbaum
arXiv:2402. 18121v2 Announce Type: replace-cross Abstract: This study assesses four cutting-edge language models in the underexplored Aminoacian language.
By Yunze Xiao, Yiyang Pan
arXiv:2608.30498v1 Announce Type: new
Abstract: Multimodal Large Language Models (MLLMs) have shown remarkable success in STEM domains, where progress is often driven by vertical, step-by-step deduct...
By Qi Li, Zhaojie Kang, Yingjie He, Zheng Lin, Hao Zhang, Guangxin Wu, Yan Gong, Rong Fu, Jianyuan Ni
arXiv:2609.38486v1 Announce Type: cross
Abstract: Large Language Models (LLMs) and more broadly Artificial Intelligence (AI) systems are often described and understood in human-like terms, a phenomen...
By Ismael T. Freire, Marceau Nahon, Maud van Lier, Katie Evans, H\'elie Bazin, Michele Farisco, Kathinka Evers, Raja Chatila, Mehdi Khamassi
arXiv:2408. 05568v2 Announce Type: replace Abstract: Large Language Models (LLMs) exhibit potentially harmful biases that reinforce culturally embedded stereotypes, influence moral judgments, or amplify positive evaluations of majority groups.
By Florian Scholten, Tobias R. Rebholz, Mandy H\"utter
arXiv:2506. 16697v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are entering psychological research both as tools and as objects of inquiry.
By Zhicheng Lin
The study evaluates eleven autoregressive transformer models on English agreement attraction scenarios using a surprisal-based approach. Results show that while transformers match human reading times for prepositional phrase configurations, they perform poorly on object‑extracted relative clauses, with predictions diverging across models and failing to capture human interference patterns. The authors argue that current transformers cannot adequately model human morphosyntactic processing and call for more rigorous, comprehensive testing to avoid misleading conclusions from limited syntactic setups.
By Titus von der Malsburg, Sebastian Pad\'o
arXiv:2509.24877v3 Announce Type: replace
Abstract: The social science of large language models (LLMs) examines how these systems evoke mind attributions, interact with one another, and transform hum...
By Xiao Jia, Zhanzhan Zhao
arXiv:2606. 05531v1 Announce Type: cross Abstract: Despite the rapid progress of Vision-Language Models (VLMs), the field lacks benchmarks that rigorously diagnose their true reasoning abilities and chart meaningful progress toward human-like multimodal intelligence.
By Mohammad Mahdi Abootorabi, Omid Ghahroodi, Anas Madkoor, Marzia Nouri, Doratossadat Dastgheib, Mohamed Hefeeda, Ehsaneddin Asgari
Multimodal Large Language Models (MLLMs) have shown remarkable success in STEM domains, where progress is often driven by vertical, step-by-step deduction under relatively stable symbol systems. Their...