Understanding the capabilities, limitations, and societal impact of large language models
Related stories
Large Language Models: A New Moore's Law?
The Reformer - Pushing the limits of language modeling
Red-Teaming Large Language Models
The Curse of Multilinguality in Lexical Normalization
The paper investigates how many languages should be jointly trained in a single lexical normalization model. Using a fixed-capacity character-level model across twelve languages, it finds that accuracy peaks when a language is trained with only a few others—typically one to four—and then declines sharply as more languages are added, dropping about forty percent. A control experiment keeping total training data constant shows the decline is due to competition for model capacity rather than data scarcity, and no reliable typological rule predicts the optimal number of co‑training languages.
Introducing The World's Largest Open Multilingual Language Model: BLOOM
Language, Language Models, and What We're Talking About
The paper discusses how language models, often treated as technical artifacts, are actually shaped by the linguistic data used in their training. Using Italian language models trained on translated and synthetic data, the author questions whether these models truly represent Italian or language more broadly, and whether NLP should focus on producing natural language. The work calls for a clearer distinction between models built as products and those built as tools for linguistic study, suggesting that diverse answers and languages may emerge without necessarily being pessimistic.
Language Models for Portuguese: A Systematic Mapping Study
The paper surveys language models created for Portuguese, noting that while rapid progress has been made in NLP, development has been uneven across languages. It systematically maps 46 Portuguese models, detailing aspects such as base model, architecture, resources, datasets, licensing, code, data, and weights. The study also traces model evolution phylogenetically, highlights research gaps, and outlines future directions for Portuguese language modeling.
Setting Up Your Own Large Language Model
Still a long way to go, but the future is promising The post Setting Up Your Own Large Language Model appeared first on Towards Data Science .
surprisal is Not a Theory
The article argues that Surprisal Theory, often presented as a computational-level explanation, is not a theory in its own right. It contends that using large language model (LLM) surprisals without considering the underlying representational and algorithmic choices obscures the theory’s commitments. The authors demonstrate through three analyses that algorithm and architecture significantly influence language model probabilities, urging researchers to reassess treating LLM surprisals as interchangeable.
Economic impacts research at OpenAI
Call for expressions of interest to study the economic impacts of large language models.
The State Of LLMs 2025: Progress, Problems, and Predictions
A 2025 review of large language models, from DeepSeek R1 and RLVR to inference-time scaling, benchmarks, architectures, and predictions for 2026.
