Large language models

Model releases, architecture work and prompting research on large language models — from frontier-lab announcements to the arXiv papers behind them.

25,737 stories · RSS feed

arXiv Computation and Language
1d ago

Common Corpus: The Largest Collection of Ethical Data for LLM Pre-Training

arXiv:2506.01732v4 Announce Type: replace Abstract: Large Language Models (LLMs) are pre-trained on large amounts of data from different sources and domains. Such datasets often contain trillions of...

By Pierre-Carl Langlais, Pavel Chizhov, Catherine Arnett, Carlos Rosas-Hinostroza, Mattia Nee, Eliot Krzystof Jones, Ir\`ene Girard, David Mach, Anastasia Stasenko, Ivan P. Yamshchikov
arXiv Computation and Language
1d ago

Inductive Claims Extraction at Scale

arXiv:2610.05275v2 Announce Type: replace Abstract: A large part of political discourse on social media is built and expressed at a level of claims: i.e. declarative, typically single-clause statemen...

By Sandrine Chausson, Bj\"orn Ross
arXiv Computer Vision
1d ago

Scaling Laws for Deepfake Detection

arXiv:2510.16320v2 Announce Type: replace Abstract: This paper presents a systematic study of scaling laws for the deepfake detection task. Specifically, we analyze the model performance against the...

By Wenhao Wang, Jusheng Zhang, Longqi Cai, Taihong Xiao, Yuxiao Wang, Ming-Hsuan Yang