Hugging Face Blog

XLSCOUT Unveils ParaEmbed 2.0: a Powerful Embedding Model Tailored for Patents and IP with Expert Support from Hugging Face

arXiv AI
Aug 19

Sparse Coverage: Semantic Center Representations for Patent Prior-Art Retrieval

Sparse Coverage is an unsupervised semantic retrieval framework designed for patent prior‑art search. It maps local span embeddings to a sparse vocabulary of embedding‑space centers chosen via a coverage‑oriented k‑center objective, allowing spans to activate nearby centers and produce sparse representations that work with inverted‑index retrieval. Experiments on CLEF‑IP 2013 demonstrate that Sparse Coverage matches or surpasses dense patent encoders in document‑level recall while remaining competitive at the passage level, making it an effective first‑stage retrieval approach for patent search.

By You Zuo (ALMAnaCH), Kim Gerdes (LISN, Qatent, STL), \'Eric de la Clergerie (ALMAnaCH), Beno\^it Sagot (ALMAnaCH)
arXiv Machine Learning
Aug 27

GreenLeaf Law Embed Tiny: A Compact Embedding Model for Legal Domain Retrieval

GreenLeaf Law Embed Tiny is a 0.6 B parameter embedding model designed for legal domain retrieval. It achieves 75.11 % on the Massive Legal Embedding Benchmark and 64.38 % on MTEB(Law, v1), outperforming other models under 1 B parameters. The model is trained via a two‑stage pipeline that distills knowledge from a larger teacher, fine‑tunes with hard negative mining, and uses a curated dataset of 3.4 million query‑passage pairs, including 150,000 human‑curated samples from diverse legal jurisdictions, while supporting efficient inference with multiple quantization levels.

By Surya Saka