arXiv AI

Counting Documents Is Not Counting Text: Unit Bias in Web-PDF Corpus Statistics

arXiv:2608. 16390v1 Announce Type: cross Abstract: PDF corpora advertise their size in tokens but compute every rate they publish (coverage, OCR routing, re-fetch recovery, language mix) per document, and none decomposes its token total.