arXiv AI
The article presents a minimal working model for large language model (LLM) systems, emphasizing four key distinctions—pretraining vs. deployment, distribution vs. samples, types of memory, and task competence vs. agency. Using this framework, it diagnoses six common misconceptions about LLMs (next‑token prediction, regression to the mean, training‑data regurgitation, model memory, alignment, and understanding), explaining what each misconception captures correctly, where it conflates distinctions, and the implications for evaluation, design, and governance. The model is applied to AI policy language, illustrating how policy can misrepresent these distinctions and offering a diagnostic toolkit to correct such errors.
arXiv:2607. 01248v1 Announce Type: cross Abstract: Large language models are increasingly used for knowledge acquisition, code generation, academic writing, and agent-based automation.
By Yang Zhao, Yingshuo Li, Zeyu Zhang
Emergent Abilities in Large Language Models: A Survey reviews how scaling LLMs leads to previously unseen capabilities such as advanced reasoning, in-context learning, coding, and problem-solving. The paper critically examines definitions, inconsistencies, and the conditions that foster these abilities, including scaling laws, task complexity, pre‑training loss, quantization, and prompting strategies. It also discusses the extension to Large Reasoning Models and highlights safety concerns like deception, manipulation, and reward hacking, calling for improved evaluation and governance.
By Leonardo Berti, Flavio Giorgi, Gjergji Kasneci
arXiv:2603. 28371v2 Announce Type: replace-cross Abstract: When an agent can articulate why something works, we typically take this as evidence of genuine understanding.
By Camilo Chac\'on Sartori
The paper defends the 'Whole Hog Thesis', arguing that sophisticated large language models such as ChatGPT are full linguistic and cognitive agents, possessing understanding, beliefs, desires, knowledge, and intentions. It rejects low‑level computational starting points and instead builds its case from high‑level behavioral observations, using Holistic Network Assumptions to link actions to mental states. The authors systematically rebut common objections—such as hallucinations and planning errors—by showing these resemble human fallibility and by challenging the necessity of traditional conditions like embodiment or semantic grounding.
By Herman Cappelen, Josh Dever
arXiv:2602. 12430v4 Announce Type: replace-cross Abstract: The transition from monolithic language models to modular, skill-equipped agents marks a defining shift in how large language models (LLMs) are deployed in practice.
By Renjun Xu, Yang Yan
arXiv:2606. 12713v1 Announce Type: new Abstract: Claims that artificial general intelligence has already arrived and claims that it remains decades away are often defended from overlapping evidence.
By J. E. Aguilera Briones
arXiv:2607. 25620v1 Announce Type: new Abstract: Quattrociocchi and colleagues warn that the fluent outputs of large language models may allow linguistic plausibility to substitute for epistemic evaluation, producing the condition they call *Epistemia*: the experience of possessing knowledge without undertaking the practices through which judgment would ordinarily be warranted.
By Federico Cabitza, Gianluca Colombo
arXiv:2607. 00048v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used in exam- and certification-style question answering tasks, where their ability to retrieve, interpret, and apply domain-specific knowledge can be systematically assessed.
By Robson Alves Vilar, Emanuel Dantas Filho, Ademar Fran\c{c}a de Sousa Neto, Mirko Perkusich, Danyllo Wagner Albuquerque, Jo\~ao Paiva, Kyller Gorg\^onio, Angelo Perkusich
arXiv:2604. 19139v3 Announce Type: replace-cross Abstract: As Large Language Models (LLMs) continue to evolve through alignment techniques such as Reinforcement Learning from Human Feedback (RLHF) and Constitutional AI, a growing and increasingly conspicuous phenomenon has emerged: the proliferation of verbal tics--repetitive, formulaic linguistic patterns that pervade model outputs.
By Shuai Wu, Xue Li, Yanna Feng, Yufang Li, Zhijun Wang, Ran Wang
arXiv:2607. 18725v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly fine-tuned for critical-domain Question-Answering (QA), yet choosing which small model to adapt, before paying the cost of adaptation, remains difficult.
By Shaswata Mitra, Subash Neupane, Trisha Chakraborty, Himanshu Tripathi, Sudip Mittal, Aritran Piplai, Shahram Rahimi
arXiv:2607. 16057v1 Announce Type: cross Abstract: Large language models (LLMs) are improving rapidly as reflected in benchmark scores, yet these AI benchmarks largely test capabilities such as factual recall, narrow question answering, mathematical problem-solving, and coding and agentic tool-use.
By Ajay Patel, Kartik Hosanagar, Ramayya Krishnan, Chris Callison-Burch, Karim Lakhani, Mitch Weiss
arXiv:2606. 06546v1 Announce Type: new Abstract: Evaluating large language models (LLMs) for education requires measuring how models teach, not only what they know.
By Tao Liu, Ye Lu, Ruohua Zhang, Siyu Song, Wentao Liu, Aimin Zhou, Hao Hao