Hugging Face Trending Papers

Scalable Frequency- and Length-Aware Subdocument Deduplication for Large Language Model Pretraining

Large-scale pretraining corpora contain substantial duplicate content. Although document-level deduplication is widely used, removing subdocument-level redundancy remains challenging.

arXiv AI
Jul 28

CuraWeb: Joint Optimization of Quality, Redundancy, and Diversity for Web-Scale Pretraining Data

arXiv:2607. 22662v1 Announce Type: new Abstract: Open-web corpora curated via highly selective filters, such as FineWeb-Edu and DCLM, constitute the core of LLM pretraining data and have significantly advanced LLM performance.

By Peiguang Li, Yongwei Zhou, Juncheng Diao, Yuchun Fan, Jian Yang, Jianxiao Yang, Zhongda Su, Shuguang Jiao, Xiao Wei, Zhiye Zou, Gan Dong, Zhizhao Zeng, Rongxiang Weng, Jingang Wang, Xunliang Cai
Hugging Face Trending Papers
Aug 19

Institutional Books - Enriched Text: A customizable multilingual open-source pipeline for denoising, deduplicating, and annotating OCR text at scale

Institutional Books – Enriched Text is a 2025 release that transforms Harvard Library’s 983,004-volume collection (IB‑HL) into a multilingual, annotated dataset. The pipeline normalizes OCR text while preserving metadata, separating endmatter, detecting paragraph language, clustering duplicates, and scoring bits‑per‑byte, all wrapped in HTML‑like annotations. The resulting IB‑HL‑ET contains 217 B tokens across 983,003 volumes and 1.39 B annotated subtopic paragraphs, enabling users to customize output rather than accept a single editorial decision.

arXiv AI
Sep 17

GVD: Governed Versioning and Deduplication for Document Repositories

GVD (Governed Versioning and Deduplication) is a framework that unifies cross‑document version linking with rule‑level conflict resolution under an auditable update policy. It assigns incoming documents to version families via bidirectional rule alignment, then compares their rules against family memory to detect duplicates, contradictions, asymmetric refinements, and new knowledge, using Counterfactual Span Probing (CSP) to correct neutral misclassifications. In a test on 120 enterprise documents across 59 version families, GVD achieved an F1 of 0.97 for version‑family construction and 0.94 for rule‑level consistency, with CSP improving rule consistency from 0.90 to 0.94.

By Mohammadreza Sediqin, Shivali Dalmia, Sumukha Thoppanahalli, Abhishek Mukherji
arXiv AI
Sep 10

Fine PT-PT Web: A High-Quality 41 Billion Tokens Data Collection of the European Portuguese Web

The paper introduces a pipeline for creating a high‑quality European Portuguese (PT‑PT) web corpus, drawing from 411 TB of raw data from Arquivo.pt. It adds a novel post‑scraping step that removes boilerplate and duplicate lines before filtering, boosting the final document yield by 19.04%. The pipeline also incorporates language identification, weighted fuzzy deduplication, and neural quality classification to produce a clean, representative dataset suitable for large‑language‑model pre‑training.

By Gon\c{c}alo Vinagre, Rui Pedro Guerra, Pedro Gomes, Miguel Moura Ramos, Duarte Miguel Alves, Afonso Simpl\'icio, Diogo Tavares, David Semedo, Daniel Gomes, Jo\~ao Magalh\~aes
arXiv AI
Aug 3

An Ontology-Guided, Deduplication-Aware Extraction Layer for Knowledge Graph Construction from Heterogeneous Documents

arXiv:2607. 28662v1 Announce Type: new Abstract: Large language models extract entities and relationships from unstructured documents fluently but inconsistently: type vocabularies fracture across documents, the same person surfaces under several name variants, relationships duplicate, and distinct individuals who share a name risk silent conflation.

By Vaibhav Dangaich, Kevin Lewis, Kundeshwar Pundalik
arXiv Machine Learning
Aug 28

Stack Trace-Based Crash Deduplication with Transformer Adaptation

Stack Trace-Based Crash Deduplication with Transformer Adaptation introduces dedupT, a transformer‑based method that models entire stack traces instead of isolated frames. The approach first fine‑tunes a pretrained language model on stack traces and then trains a fully‑connected network to rank duplicate crashes. Experiments on four public datasets show dedupT improves Mean Reciprocal Rank by over 15% versus the best deep‑learning baseline and up to 10% over traditional methods, while also achieving higher ROC‑AUC for unique crash detection.

By Md Afif Al Mamun, Gias Uddin, Lan Xia, Longyu Zhang