arXiv AI By Lars Henry Berge Olsen, Pierre Lison, Martin Jullum, Mark Anderson

FindMyText: Robust, Scalable Detection of Text Containment in Large Web-Crawled Corpora

Read the original on arXiv AI →

arXiv:2607. 10020v2 Announce Type: replace-cross Abstract: We present FindMyText, an open-source Python package designed to efficiently assess whether a given text appears, in part or in full, within a text corpus.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Aug 19

Institutional Books - Enriched Text: A customizable multilingual open-source pipeline for denoising, deduplicating, and annotating OCR text at scale

Institutional Books – Enriched Text is a 2025 release that transforms Harvard Library’s 983,004-volume collection (IB‑HL) into a multilingual, annotated dataset. The pipeline normalizes OCR text while preserving metadata, separating endmatter, detecting paragraph language, clustering duplicates, and scoring bits‑per‑byte, all wrapped in HTML‑like annotations. The resulting IB‑HL‑ET contains 217 B tokens across 983,003 volumes and 1.39 B annotated subtopic paragraphs, enabling users to customize output rather than accept a single editorial decision.

arXiv AI
Jun 9

Efficient and Scalable Provenance Tracking for LLM-Generated Code Snippets

arXiv:2605. 28510v2 Announce Type: replace-cross Abstract: Large language models (LLMs) for code completion and generation are increasingly used in software development, yet they may reproduce training examples verbatim and without authorship attribution, raising legal and ethical concerns around plagiarism and license compliance.

By Andrea Gurioli, Davide D'Ascenzo, Federico Pennino, Maurizio Gabbrielli, Stefano Zacchiroli