arXiv AI By Luca Foppiano

Counting Documents Is Not Counting Text: Unit Bias in Web-PDF Corpus Statistics

Read the original on arXiv AI →

arXiv:2608. 16390v1 Announce Type: cross Abstract: PDF corpora advertise their size in tokens but compute every rate they publish (coverage, OCR routing, re-fetch recovery, language mix) per document, and none decomposes its token total.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.