arXiv AI

Six misconceptions about large language models: A minimal model and diagnostic taxonomy

The article presents a minimal working model for large language model (LLM) systems, emphasizing four key distinctions—pretraining vs. deployment, distribution vs. samples, types of memory, and task competence vs. agency. Using this framework, it diagnoses six common misconceptions about LLMs (next‑token prediction, regression to the mean, training‑data regurgitation, model memory, alignment, and understanding), explaining what each misconception captures correctly, where it conflates distinctions, and the implications for evaluation, design, and governance. The model is applied to AI policy language, illustrating how policy can misrepresent these distinctions and offering a diagnostic toolkit to correct such errors.

arXiv Machine Learning
1d ago

Emergent Abilities in Large Language Models: A Survey

Emergent Abilities in Large Language Models: A Survey reviews how scaling LLMs leads to previously unseen capabilities such as advanced reasoning, in-context learning, coding, and problem-solving. The paper critically examines definitions, inconsistencies, and the conditions that foster these abilities, including scaling laws, task complexity, pre‑training loss, quantization, and prompting strategies. It also discusses the extension to Large Reasoning Models and highlights safety concerns like deception, manipulation, and reward hacking, calling for improved evaluation and governance.

By Leonardo Berti, Flavio Giorgi, Gjergji Kasneci
arXiv AI
4d ago

Recognizing Artificial Minds: A Philosophical Defense of AI Cognition

The paper defends the 'Whole Hog Thesis', arguing that sophisticated large language models such as ChatGPT are full linguistic and cognitive agents, possessing understanding, beliefs, desires, knowledge, and intentions. It rejects low‑level computational starting points and instead builds its case from high‑level behavioral observations, using Holistic Network Assumptions to link actions to mental states. The authors systematically rebut common objections—such as hallucinations and planning errors—by showing these resemble human fallibility and by challenging the necessity of traditional conditions like embodiment or semantic grounding.

By Herman Cappelen, Josh Dever
arXiv AI
Jul 29

Beyond Epistemia: Epistemic Schizologia and Large Language Models as Techno-Semiotic Machines

arXiv:2607. 25620v1 Announce Type: new Abstract: Quattrociocchi and colleagues warn that the fluent outputs of large language models may allow linguistic plausibility to substitute for epistemic evaluation, producing the condition they call *Epistemia*: the experience of possessing knowledge without undertaking the practices through which judgment would ordinarily be warranted.

By Federico Cabitza, Gianluca Colombo
arXiv AI
Jul 2

Comparing Large Language Models on Scrum Certification-Style Questions: Accuracy, Stability, and Error Patterns

arXiv:2607. 00048v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used in exam- and certification-style question answering tasks, where their ability to retrieve, interpret, and apply domain-specific knowledge can be systematically assessed.

By Robson Alves Vilar, Emanuel Dantas Filho, Ademar Fran\c{c}a de Sousa Neto, Mirko Perkusich, Danyllo Wagner Albuquerque, Jo\~ao Paiva, Kyller Gorg\^onio, Angelo Perkusich
arXiv AI
Jul 7

The Rise of Verbal Tics in Large Language Models: A Systematic Analysis Across Frontier Models

arXiv:2604. 19139v3 Announce Type: replace-cross Abstract: As Large Language Models (LLMs) continue to evolve through alignment techniques such as Reinforcement Learning from Human Feedback (RLHF) and Constitutional AI, a growing and increasingly conspicuous phenomenon has emerged: the proliferation of verbal tics--repetitive, formulaic linguistic patterns that pervade model outputs.

By Shuai Wu, Xue Li, Yanna Feng, Yufang Li, Zhijun Wang, Ran Wang
arXiv AI
Jul 22

Find Before You Fine-Tune: A Diagnostic Study of Small LLMs for Cybersecurity QA

arXiv:2607. 18725v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly fine-tuned for critical-domain Question-Answering (QA), yet choosing which small model to adapt, before paying the cost of adaptation, remains difficult.

By Shaswata Mitra, Subash Neupane, Trisha Chakraborty, Himanshu Tripathi, Sudip Mittal, Aritran Piplai, Shahram Rahimi
arXiv AI
Jul 20

Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning

arXiv:2607. 16057v1 Announce Type: cross Abstract: Large language models (LLMs) are improving rapidly as reflected in benchmark scores, yet these AI benchmarks largely test capabilities such as factual recall, narrow question answering, mathematical problem-solving, and coding and agentic tool-use.

By Ajay Patel, Kartik Hosanagar, Ramayya Krishnan, Chris Callison-Burch, Karim Lakhani, Mitch Weiss