arXiv Machine Learning

Can Subgraph Explanations Be Weaponized to Steal Graph Neural Networks?

Graph Machine Learning as a Service platforms now offer explainability interfaces to satisfy regulatory transparency, but this transparency can be exploited. The paper introduces a novel model extraction attack for graph classification that operates under strict black‑box constraints, using only discrete class labels and binary explanation masks. The method guides Monte Carlo edge sensitivity estimation toward decision boundaries with Hoeffding guarantees and narrows the search space using explanation subgraphs, outperforming comparable baselines on benchmark datasets.