The Reformer - Pushing the limits of language modeling
Related stories
surprisal is Not a Theory
The article argues that Surprisal Theory, often presented as a computational-level explanation, is not a theory in its own right. It contends that using large language model (LLM) surprisals without considering the underlying representational and algorithmic choices obscures the theory’s commitments. The authors demonstrate through three analyses that algorithm and architecture significantly influence language model probabilities, urging researchers to reassess treating LLM surprisals as interchangeable.
Aligning language models to follow instructions
Scaling laws for neural language models
How to train a new language model from scratch using Transformers and Tokenizers
Red-Teaming Large Language Models
Very Large Language Models and How to Evaluate Them
Large Language Models: A New Moore's Law?
Language, Language Models, and What We're Talking About
The paper discusses how language models, often treated as technical artifacts, are actually shaped by the linguistic data used in their training. Using Italian language models trained on translated and synthetic data, the author questions whether these models truly represent Italian or language more broadly, and whether NLP should focus on producing natural language. The work calls for a clearer distinction between models built as products and those built as tools for linguistic study, suggesting that diverse answers and languages may emerge without necessarily being pessimistic.
The Grammar of Transformers: A Systematic Review of Interpretability Research on Syntactic Knowledge in Language Models
arXiv:2601.19926v3 Announce Type: replace-cross Abstract: We present a systematic review of 337 articles evaluating the syntactic abilities of Transformer-based language models (TLMs), reporting on o...
Accelerating Vision-Language Models: BridgeTower on Habana Gaudi2
Language Models for Portuguese: A Systematic Mapping Study
The paper surveys language models created for Portuguese, noting that while rapid progress has been made in NLP, development has been uneven across languages. It systematically maps 46 Portuguese models, detailing aspects such as base model, architecture, resources, datasets, licensing, code, data, and weights. The study also traces model evolution phylogenetically, highlights research gaps, and outlines future directions for Portuguese language modeling.