Private Seeds, Public LLMs: Realistic and Privacy-Preserving Synthetic Data Generation
arXiv:2604. 07486v3 Announce Type: replace-cross Abstract: Large language models (LLMs) have emerged as a powerful tool for synthetic data generation.
arXiv:2601. 04641v2 Announce Type: replace-cross Abstract: The deployment of Machine-Generated Text (MGT) detection systems necessitates processing sensitive user data, creating a fundamental conflict between authorship verification and privacy preservation.
arXiv:2604. 07486v3 Announce Type: replace-cross Abstract: Large language models (LLMs) have emerged as a powerful tool for synthetic data generation.
The paper introduces a style-aware paraphrasing method for text anonymization that leverages pretrained large language models to build compact stylistic profiles from minimal samples and rewrite text to suppress identifiable style markers while preserving meaning. It demonstrates that this approach reduces authorship attribution F1 scores by 60‑70% on blog and review datasets, outperforming both differential privacy‑based and non‑DP baselines, and maintains content quality and readability.
DP-IPI introduces a hybrid differential privacy text rewriting mechanism that selectively privatizes only the spans containing indirect personal identifiers (IPIs) in clinical texts. By targeting these specific tokens rather than all words, the method preserves higher text quality and usability while still reducing re-identification risks. The approach demonstrates improved privacy‑utility trade‑offs compared to indiscriminate DP text rewriting techniques.
arXiv:2608.29624v1 Announce Type: new Abstract: Natural Language Processing methods have enabled novel solutions and advances in the field of privacy, particularly in the sub-domain of text-to-text p...
The paper introduces the privacy‑HSD trade‑off, highlighting that automatic hate speech detection systems can inadvertently compromise user privacy by encoding authorship. It demonstrates that such systems may achieve high performance at the expense of privacy, and proposes a new domain‑specific technique, AgnoSpeech, alongside other text privatization methods to balance these competing goals. The authors benchmark these methods, showing that while challenging, it is feasible to protect privacy without sacrificing hate‑speech detection effectiveness.
Redakto is a new tool designed to anonymize text before it is processed by large language models (LLMs). It offers state‑of‑the‑art redaction of personally identifiable information (PII) and pseudonymization, accessible via a web interface, REST APIs, and model context protocol hooks. The authors evaluate its performance on legal and medical datasets, showing that anonymized texts retain utility comparable to the originals, enabling LLM tasks without significant loss of effectiveness.
arXiv:2606. 24408v1 Announce Type: new Abstract: Assessing the privacy of large language models (LLMs) presents significant challenges.
QuanText is a training‑free, large‑language‑model‑agnostic mechanism for releasing textual datasets that protects dataset‑level secrets such as the proportion of records with a particular diagnosis or gender. It perturbs both the secret distribution and correlated attribute distributions by selecting candidate release distributions close to the private empirical distribution and rewriting each text sample to match the chosen distribution using attribute‑related snippets. The method is inspired by the Statistic Maximal Leakage framework and, under idealized conditions, satisfies an SML guarantee, while empirical evaluations show a superior privacy‑utility trade‑off compared to existing data generation baselines.
arXiv:2609.38934v1 Announce Type: cross Abstract: Differentially private (DP) text generation can protect individual records, but privacy alone does not specify what evidence a released statement car...
arXiv:2606. 24623v1 Announce Type: cross Abstract: Retrieval-Augmented Generation enhances large language models by incorporating external knowledge, but deploying it in sensitive scenarios risks privacy leakage via malicious prompts.
arXiv:2606. 16952v2 Announce Type: replace-cross Abstract: The rapid adoption of generative AI and Large Language Models (LLMs) has spurred interest in synthetic data as a privacy-preserving alternative to sensitive real-world datasets.
The paper systematically studies how anonymizing input data affects large language models (LLMs). Five prominent LLMs were evaluated on eleven benchmarks, comparing performance on original versus pseudonymized inputs. Results show that anonymization generally degrades performance, with larger drops for more capable models and task-dependent effects; reversible anonymization preserves entity uniqueness better than irreversible redaction, and prompting about anonymization offers no benefit.