arXiv AI By Lars Henry Berge Olsen, Pierre Lison, Martin Jullum, Mark Anderson

FindMyText: Robust, Scalable Detection of Text Containment in Large Web-Crawled Corpora

Read the original on arXiv AI →

arXiv:2607. 10020v2 Announce Type: replace-cross Abstract: We present FindMyText, an open-source Python package designed to efficiently assess whether a given text appears, in part or in full, within a text corpus.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.

arXiv AI
Jun 9

Efficient and Scalable Provenance Tracking for LLM-Generated Code Snippets

arXiv:2605. 28510v2 Announce Type: replace-cross Abstract: Large language models (LLMs) for code completion and generation are increasingly used in software development, yet they may reproduce training examples verbatim and without authorship attribution, raising legal and ethical concerns around plagiarism and license compliance.

By Andrea Gurioli, Davide D'Ascenzo, Federico Pennino, Maurizio Gabbrielli, Stefano Zacchiroli