Large scale document deduplication must preserve semantic equivalence while remaining efficient over massive corpora. We present SemHash LLM, a multi granularity framework that unifies semantic projection hashing, attention weighted MinHash, contrastive boundary learning, and selective LLM based adjudication.
arXiv:2607. 01601v1 Announce Type: new Abstract: Large scale document deduplication must preserve semantic equivalence while remaining efficient over massive corpora.
By Xinyi Fang, Kejian Tong, Jiabei Liu, Tao Ning, Yuhang He
arXiv:2607. 22662v1 Announce Type: new Abstract: Open-web corpora curated via highly selective filters, such as FineWeb-Edu and DCLM, constitute the core of LLM pretraining data and have significantly advanced LLM performance.
By Peiguang Li, Yongwei Zhou, Juncheng Diao, Yuchun Fan, Jian Yang, Jianxiao Yang, Zhongda Su, Shuguang Jiao, Xiao Wei, Zhiye Zou, Gan Dong, Zhizhao Zeng, Rongxiang Weng, Jingang Wang, Xunliang Cai
Institutional Books – Enriched Text is a 2025 release that transforms Harvard Library’s 983,004-volume collection (IB‑HL) into a multilingual, annotated dataset. The pipeline normalizes OCR text while preserving metadata, separating endmatter, detecting paragraph language, clustering duplicates, and scoring bits‑per‑byte, all wrapped in HTML‑like annotations. The resulting IB‑HL‑ET contains 217 B tokens across 983,003 volumes and 1.39 B annotated subtopic paragraphs, enabling users to customize output rather than accept a single editorial decision.
arXiv:2608.28622v1 Announce Type: cross
Abstract: Large language model training corpora grow through successive, often redundant releases, so each release must be deduplicated against both itself and...
By Xiao Yang, Erik Edward Aldape, Beren Millidge
GVD (Governed Versioning and Deduplication) is a framework that unifies cross‑document version linking with rule‑level conflict resolution under an auditable update policy. It assigns incoming documents to version families via bidirectional rule alignment, then compares their rules against family memory to detect duplicates, contradictions, asymmetric refinements, and new knowledge, using Counterfactual Span Probing (CSP) to correct neutral misclassifications. In a test on 120 enterprise documents across 59 version families, GVD achieved an F1 of 0.97 for version‑family construction and 0.94 for rule‑level consistency, with CSP improving rule consistency from 0.90 to 0.94.
By Mohammadreza Sediqin, Shivali Dalmia, Sumukha Thoppanahalli, Abhishek Mukherji