Emergent Abilities in Large Language Models: A Survey reviews how scaling LLMs leads to previously unseen capabilities such as advanced reasoning, in-context learning, coding, and problem-solving. The paper critically examines definitions, inconsistencies, and the conditions that foster these abilities, including scaling laws, task complexity, pre‑training loss, quantization, and prompting strategies. It also discusses the extension to Large Reasoning Models and highlights safety concerns like deception, manipulation, and reward hacking, calling for improved evaluation and governance.
By Leonardo Berti, Flavio Giorgi, Gjergji Kasneci
arXiv:2606. 20231v1 Announce Type: new Abstract: Can intelligence be measured?
By Ishanu Chattopadhyay
arXiv:2606. 02632v1 Announce Type: cross Abstract: Modern Machine Learning (ML) and Artificial Intelligence (AI) models, especially large language models (LLMs), are increasingly used to generate scientific hypotheses and mechanistic explanations from observational data.
By Tyler H. McCormick
arXiv:2607. 15883v1 Announce Type: cross Abstract: Large language models are broadly capable, yet in sustained one-to-one conversation they still read as flat: competent, responsive, and somehow not quite the presence of a mind.
By Sebastian Cochinescu
The paper introduces a framework that combines world models, which generate concrete visual rollouts of possible futures, with multimodal large language models (MLLMs) that perform abstract reasoning. It proposes a controlled concrete reasoning approach and a new training method called Privileged‑Future On‑Policy Self‑Distillation (PF‑OPSD), which uses ground‑truth future videos as privileged teacher context during training while the student model never sees true futures at test time. Experiments on two human‑verified benchmarks, VRQABench and OpenWorldQA, show that PF‑OPSD improves performance by about 10–11% over baselines and enhances robustness to noisy or conflicting rollouts.
By Yucheng Zhou, Wei Tao, Yiwen Guo, Jianbing Shen
arXiv:2606. 17289v1 Announce Type: new Abstract: AI systems based on artificial neural networks are being developed with aspirations of pushing the boundary of human mathematical knowledge.
By Phoebe Zeng, Thomas L. Griffiths, Brenden M. Lake
arXiv:2607. 00627v1 Announce Type: new Abstract: Large language models (LLMs) are powerful pattern-completion systems, but their default operating mode - predicting the next token from a static context - does not reliably produce persistent, manipulable representations of an external world.
By Alexey Potapov
The article reports that large language models can predict and collaboratively modulate human memory search during a semantic fluency task. By tracking and forecasting participants’ semantic retrieval patterns, the models outperform other humans in following these mental trajectories. This suggests that AI can serve as a cognitive tool to extend human abilities in open‑ended conceptual exploration and creative ideation.
By Eric Lacosse, Mariana Duarte, Graham Todd, Peter M. Todd, Daniel C. McNamee
arXiv:2608. 11215v1 Announce Type: new Abstract: Simulating societies of many large language model (LLM) agents is expensive, yet the questions asked of such simulations are usually macroscopic: phase behaviour, stylised facts, and scaling with the number of agents $N$, not the cognition of any single agent.
By Igor Itkin
arXiv:2603.18007v2 Announce Type: replace-cross
Abstract: The study explores whether current Large Language Models (LLMs) exhibit Theory of Mind (ToM) capabilities -- specifically, the ability to inf...
By Anna Babarczy, Andras Lukacs, Peter Vedres, Zeteny Bujka
arXiv:2608. 16213v1 Announce Type: new Abstract: Intelligence is constituted by \textit{process} (iterative activity through which output emerges), not in the output itself.
By Michael J. Richardson, Ayeh Alhasan, Cassandra Crone, M. Paula Diaz Monfort, Patrick Nalepka, Mark Dras, Rachel W. Kallen, David M. Kaplan
CogGym is a scalable, unified framework that standardizes diverse cognitive experiments into a task‑agnostic Experiment Markup Language (EML) for systematic comparison of human and AI behavior. The initial release curates 258 experiments from 100 papers focused on human commonsense reasoning and evaluates 50 large language models, revealing a scaling trend where larger models better reproduce human judgments but still lag far behind human split‑half reliability. The framework aims to continually incorporate new cognitive science experiments to track where model behavior aligns with or diverges from human cognition as models evolve.
By Lance Ying, Jinzhou Wu, Yingshan Susan Wang, Shivam Aarya, Luca M. Schulze Buschoff, Harry Chen, Katherine M. Collins, Andrea de Varda, Shuhao Fu, Sean Dae Houlihan, Akshay K. Jagadish, Guangyuan Jiang, Samuel Kiegeland, Tetsu Kurumisawa, Rongzhi Liu, Ryan Liu, Ningshan Ma, Kathryn McGregor, Younes Strittmatter, Polina Tsvilodub, Jacob Hoover Vigly, Sarah Wu, Enjie Xu, Yiling Yun, Kelsey Allen, Tyler Brooke-Wilson, Brian Christian, Evelina Fedorenko, Michael C. Frank, Michael Franke, Tao Gao, Samuel J. Gershman, Robert D. Hawkins, Jennifer Hu, Julian Jara-Ettinger, Max Kleiman-Weiner, Sydney Levine, Tal Linzen, Hongjing Lu, Timothy O'Donnell, Desmond C. Ong, Steven T. Piantadosi, Rebecca Saxe, Eric Schulz, Tianmin Shu, Felix A. Sosa, Ilia Sucholutsky, Tan Zhi-Xuan, Tomer Ullman, Fei Xu, Ilker Yildirim, Jian-Qiao Zhu, Thomas L. Griffiths, Tobias Gerstenberg, Kevin Smith, Joshua B. Tenenbaum