arXiv Computation and Language By Ahmed Sohair Khan, Estrid He, Chenglong Ma, Monica Wachowicz, Elham Naghizade

GraphProfiler: Source-Linked Sensitive Attribute Inference via Personal Knowledge Graphs

Read the original on arXiv Computation and Language →

GraphProfiler is an auditable LLM-based profiler that infers sensitive attributes such as age, income, and occupation from user-generated content by aggregating indirect cues across many posts. It represents each user's post history as a source-linked personal knowledge graph, allowing predictions to be traced back to specific posts, concepts, and relationships. The system achieves an 86.7% success rate on the SynthPAI benchmark and 84.6% on PANDORA, citing supporting evidence for over 98% of predictions and demonstrating that cited posts significantly contribute to attack success.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Machine Learning
Aug 27

Are LLM-Enhanced GNNs Privacy-Safe?

The paper evaluates privacy risks in graph neural networks enhanced by large language models (LLMs). Using a five‑stage framework, the authors test six real‑world text‑attributed graph datasets with 42 model configurations and six privacy attack methods across link, label, and membership inference threats. Results show that LLM‑enhanced GNNs are more vulnerable than shallow baselines, with semantic enrichment amplifying exploitable signals, and that differential privacy can reduce risk but at a significant cost to utility.

By Longzhu He, Zelang Wen, Chaozhuo Li, Sen Su
arXiv Machine Learning
Sep 17

QuanText: Protecting Dataset-Level Secrets in Textual Data Sharing

QuanText is a training‑free, large‑language‑model‑agnostic mechanism for releasing textual datasets that protects dataset‑level secrets such as the proportion of records with a particular diagnosis or gender. It perturbs both the secret distribution and correlated attribute distributions by selecting candidate release distributions close to the private empirical distribution and rewriting each text sample to match the chosen distribution using attribute‑related snippets. The method is inspired by the Statistic Maximal Leakage framework and, under idealized conditions, satisfies an SML guarantee, while empirical evaluations show a superior privacy‑utility trade‑off compared to existing data generation baselines.

By Shuaiqi Wang, Zinan Lin, Giulia Fanti