← Back to all news
Hugging Face Blog May 16, 2023

Large-scale Near-deduplication Behind BigCode

Read the original on Hugging Face Blog →

The Flow has not summarised this story yet — read it at Hugging Face Blog.

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

Hugging Face Trending Papers
Aug 4

Scalable Frequency- and Length-Aware Subdocument Deduplication for Large Language Model Pretraining

Large-scale pretraining corpora contain substantial duplicate content. Although document-level deduplication is widely used, removing subdocument-level redundancy remains challenging.

llmscomputer-vision
More like this →
Hugging Face Trending Papers
Jul 2

SemHash-LLM: A Multi-Granularity Semantic Hashing Framework for Document Deduplication

Large scale document deduplication must preserve semantic equivalence while remaining efficient over massive corpora. We present SemHash LLM, a multi granularity framework that unifies semantic projection hashing, attention weighted MinHash, contrastive boundary learning, and selective LLM based adjudication.

llmsrag
More like this →
arXiv AI
Jul 3

SemHash-LLM: A Multi-Granularity Semantic Hashing Framework for Document Deduplication

arXiv:2607. 01601v1 Announce Type: new Abstract: Large scale document deduplication must preserve semantic equivalence while remaining efficient over massive corpora.

By Xinyi Fang, Kejian Tong, Jiabei Liu, Tao Ning, Yuhang He
llmsrag
More like this →
arXiv Machine Learning
Jun 15

MUFFLe: Efficient Model Update Compression via Generalized Deduplication for Federated Learning

arXiv:2606. 14354v1 Announce Type: new Abstract: Federated learning is well suited to edge environments but is often limited by the uplink cost of transmitting model updates.

By Xiaobo Zhao, Daniel E. Lucani
efficiency
More like this →
arXiv Machine Learning
Jul 28

Cross-Attention Calibrated Deduplication for Retrieval-Augmented Generation System

arXiv:2607. 24332v1 Announce Type: cross Abstract: Common chunking strategies in Retrieval-Augmented Generation (RAG) systems often create redundant chunks.

By Phuong Le Huy, Nam H. Nguyen, Quan V. Dang
rag
More like this →
Hugging Face Trending Papers
Jul 27

Where Quality Breaks in Compressed Short-Text Generation: Staged Bottleneck Localization

Compressed short-text generators can fail in two different places: the codec may discard information before generation starts, or the latent generator may produce weak codes. Without separating these failure modes, researchers can spend compute improving the wrong component.

llmsdiffusion
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea