Hugging Face Trending Papers

SemHash-LLM: A Multi-Granularity Semantic Hashing Framework for Document Deduplication

Read the original on Hugging Face Trending Papers →

Large scale document deduplication must preserve semantic equivalence while remaining efficient over massive corpora. We present SemHash LLM, a multi granularity framework that unifies semantic projection hashing, attention weighted MinHash, contrastive boundary learning, and selective LLM based adjudication.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv AI
Jul 28

CuraWeb: Joint Optimization of Quality, Redundancy, and Diversity for Web-Scale Pretraining Data

arXiv:2607. 22662v1 Announce Type: new Abstract: Open-web corpora curated via highly selective filters, such as FineWeb-Edu and DCLM, constitute the core of LLM pretraining data and have significantly advanced LLM performance.

By Peiguang Li, Yongwei Zhou, Juncheng Diao, Yuchun Fan, Jian Yang, Jianxiao Yang, Zhongda Su, Shuguang Jiao, Xiao Wei, Zhiye Zou, Gan Dong, Zhizhao Zeng, Rongxiang Weng, Jingang Wang, Xunliang Cai
arXiv AI
Sep 10

Fine PT-PT Web: A High-Quality 41 Billion Tokens Data Collection of the European Portuguese Web

The paper introduces a pipeline for creating a high‑quality European Portuguese (PT‑PT) web corpus, drawing from 411 TB of raw data from Arquivo.pt. It adds a novel post‑scraping step that removes boilerplate and duplicate lines before filtering, boosting the final document yield by 19.04%. The pipeline also incorporates language identification, weighted fuzzy deduplication, and neural quality classification to produce a clean, representative dataset suitable for large‑language‑model pre‑training.

By Gon\c{c}alo Vinagre, Rui Pedro Guerra, Pedro Gomes, Miguel Moura Ramos, Duarte Miguel Alves, Afonso Simpl\'icio, Diogo Tavares, David Semedo, Daniel Gomes, Jo\~ao Magalh\~aes
arXiv AI
2d ago

HHR: Hierarchical Hash Retrieval for Efficient LLM Generation

The paper introduces Hierarchical Hash Retrieval (HHR), a coarse‑to‑fine framework designed to improve hash‑based retrieval for large language models. HHR combines Geometry‑Aware Key Routing (GKR) to redistribute feature magnitudes and prune low‑logit keys, with Learned Hash Projection (LHP) to align Hamming distance with true query‑key relevance for fine‑grained retrieval. Experiments on diverse LLMs and benchmarks show that HHR outperforms existing methods, boosting LongBench scores by 1.10 points and achieving up to 3.30× decoding speedup at 128K context length for Llama‑3.1‑8B‑Instruct.

By Lianjun Liu, Tiantian Zheng, You Huang, Weiqi Yan, Mingte Qiu, Huazhong Liu, Xiaofeng Zhu, Yunshan Zhong