arXiv:2609.38934v1 Announce Type: cross
Abstract: Differentially private (DP) text generation can protect individual records, but privacy alone does not specify what evidence a released statement car...
By Tsubasa Takahashi, Takumi Hiraoka
arXiv:2606. 24408v1 Announce Type: new Abstract: Assessing the privacy of large language models (LLMs) presents significant challenges.
By Lorenzo Rossi, Bart{\l}omiej Marek, Franziska Boenisch, Adam Dziedzic
arXiv:2609.36326v1 Announce Type: new
Abstract: Retrieval-augmented generation (RAG) hands the user's query to whoever hosts the corpus. We propose PILLAR, a Privacy-Preserving RAG (PPRAG) system bas...
By Truong Son Nguyen (Arizona State University), Daniel Blackley (George Mason University), Ni Trieu (Arizona State University), Evgenios M. Kornaropoulos (George Mason University)
QuanText is a training‑free, large‑language‑model‑agnostic mechanism for releasing textual datasets that protects dataset‑level secrets such as the proportion of records with a particular diagnosis or gender. It perturbs both the secret distribution and correlated attribute distributions by selecting candidate release distributions close to the private empirical distribution and rewriting each text sample to match the chosen distribution using attribute‑related snippets. The method is inspired by the Statistic Maximal Leakage framework and, under idealized conditions, satisfies an SML guarantee, while empirical evaluations show a superior privacy‑utility trade‑off compared to existing data generation baselines.
By Shuaiqi Wang, Zinan Lin, Giulia Fanti
The paper introduces GROUND, a framework that limits large language model (LLM) analytics to a governed semantic layer for enterprise data warehouses. GROUND supplies approved metrics, dimensions, join paths, filters, and security rules, then validates generated SQL against these constraints before execution, retrying or abstaining on violations. In benchmarks, GROUND eliminates hallucinations across all evaluated categories and prevents row‑level security breaches, outperforming schema‑only, schema‑RAG, and semantic‑only approaches.
By Aravind Sasidharan Pillai
arXiv:2601. 11219v3 Announce Type: replace-cross Abstract: Federated learning (FL) for large language models (LLMs) has attracted increasing attention as a privacy-preserving approach for adapting models over distributed data, where parameter-efficient methods such as Low-Rank Adaptation (LoRA) are widely adopted to reduce communication and memory costs.
By Zhikang Shen, Jianrong Lu, Haiyuan Wan, Jianhai Chen
arXiv:2510. 04902v3 Announce Type: replace Abstract: Tuning hyperparameters in federated machine learning can substantially impact model performance.
By Johannes Liebenow, Thorsten Peinemann, Esfandiar Mohammadi
The paper introduces a federated inference framework that enables multiple commercial large language model (LLM) APIs—such as LLaMA‑3.3‑70B, GPT‑4o‑mini, and Claude‑3‑Haiku—to collaborate on cognitive diagnosis tasks without accessing raw student data or proprietary model internals. Each entity’s predictions are perturbed with Laplace noise to provide epsilon‑local differential privacy, and a residual‑based aggregation scheme mitigates model heterogeneity. Experiments on three educational benchmarks demonstrate strong privacy guarantees with minimal accuracy loss, confirming the framework’s practical usability and cross‑domain generalizability.
By Yagna Manasa Boyapati, Chong Yu, Tianyu Jiang, Justin Zhan
Heterogeneous Differential Privacy (HDP) in Federated Learning (FL) allows clients to select individual privacy budgets ($\varepsilon_i$) according to institutional policies and data sensitivity. In practice, many HDP-FL systems employ $\varepsilon$-aware server aggregation to improve model utility by re-weighting client updates according to their declared privacy budgets.
arXiv:2608. 00144v2 Announce Type: replace Abstract: Membership inference (MIA) on language models is usually summarised by aggregate ROC-AUC, but such evaluations are confounded: model-free blind baselines can separate members from non-members using surface text alone.
By Victor Maricato
arXiv:2608.30141v1 Announce Type: cross
Abstract: Preference optimization is widely used to align large language models with human preferences, but preference-data composition may also influence priv...
By Dishu Yang, Jingjing Liu, Jize Li
arXiv:2606. 16952v2 Announce Type: replace-cross Abstract: The rapid adoption of generative AI and Large Language Models (LLMs) has spurred interest in synthetic data as a privacy-preserving alternative to sensitive real-world datasets.
By Kareem Amin, Rudrajit Das, Alessandro Epasto, Adel Javanmard, Dennis Kraft, M\'onica Ribero, Sergei Vassilvitskii