arXiv Machine Learning By Lorenzo Rossi, Bart{\l}omiej Marek, Franziska Boenisch, Adam Dziedzic

Natural Identifiers for Privacy and Data Audits in Large Language Models

Read the original on arXiv Machine Learning →

arXiv:2606. 24408v1 Announce Type: new Abstract: Assessing the privacy of large language models (LLMs) presents significant challenges.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 17

QuanText: Protecting Dataset-Level Secrets in Textual Data Sharing

QuanText is a training‑free, large‑language‑model‑agnostic mechanism for releasing textual datasets that protects dataset‑level secrets such as the proportion of records with a particular diagnosis or gender. It perturbs both the secret distribution and correlated attribute distributions by selecting candidate release distributions close to the private empirical distribution and rewriting each text sample to match the chosen distribution using attribute‑related snippets. The method is inspired by the Statistic Maximal Leakage framework and, under idealized conditions, satisfies an SML guarantee, while empirical evaluations show a superior privacy‑utility trade‑off compared to existing data generation baselines.

By Shuaiqi Wang, Zinan Lin, Giulia Fanti