arXiv Machine Learning By William Novak (Minot State University), Muhammad Abusaqer (Minot State University)

Empirical Evaluation of Membership Inference Attacks on NLP Text Classifiers: A Baseline Study on SST-2

Read the original on arXiv Machine Learning →

The paper evaluates membership inference attacks (MIAs) on NLP text classifiers using the GLUE SST‑2 sentiment dataset. It compares a TF‑IDF + Logistic Regression pipeline with a fine‑tuned DistilBERT model under a loss‑threshold MIA, finding that both models leak membership signals despite high accuracy. The study also tests mitigations, showing that stronger regularization reduces leakage for Logistic Regression at a utility cost, while fine‑tuning DistilBERT for fewer epochs lowers leakage with minimal accuracy loss.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Aug 3

TextCloak: Thwarting Unauthorized LLM Exploitation via RL-Driven Unlearnable Text

arXiv:2607. 28862v1 Announce Type: cross Abstract: The rapid development of Large Language Models (LLMs) has led to significant advances across a wide range of language tasks, while simultaneously raising growing concerns about unauthorized data exploitation and privacy leakage.

By Chengshuai Zhao, Pingchuan Ma, Dawei Li, Bohan Jiang, Zhiyuan Yu, Zhen Tan, Huan Liu
arXiv Computation and Language
Sep 21

Conformal Privacy Auditing: Calibrated Re-identification Attacks with Statistical Guarantees

Conformal Privacy Auditing (CPA) is a distribution‑free framework that calibrates re‑identification risk for each released document against large language model (LLM)‑empowered adversaries. It outputs a conformal ambiguity set of candidate identities that is guaranteed to contain the true identity with a user‑chosen confidence level under exchangeability, along with an interpretable leakage proxy derived from the set size. CPA supports both logit‑access and sampling‑only attackers, enabling audits of both open‑source and proprietary models, and demonstrates calibrated coverage across various benchmarks and attacker configurations.

By Shuo Huang, Gholamreza Haffari, Xingliang Yuan, Ting Yu, Lizhen Qu
arXiv Machine Learning
Sep 11

Predicting Privacy Leakage from Weight Spectral Density

The paper investigates whether inexpensive spectral metrics from the heavy‑tailed self‑regularisation framework can predict membership inference attack (MIA) vulnerability, offering a scalable alternative to costly shadow‑model attacks. Experiments on image and tabular classification tasks show that stable rank correlates positively with overall MIA success, while Log alpha‑Norm correlates negatively with MIA risk in low false‑positive regimes, outperforming conventional generalisation gap measures. These findings suggest that neural network spectra contain privacy leakage signals not captured by traditional overfitting metrics, pointing to spectral analysis as a promising direction for privacy auditing.

By Richard J. Preen, Jim Smith