arXiv Machine Learning

Auditing Privacy in Multi-Tenant RAG under Account Collusion

arXiv:2605. 19847v2 Announce Type: replace-cross Abstract: Multi-tenant RAG services often treat the account as the privacy boundary: each account receives an $(\varepsilon_{\text{acc}},\delta_{\text{acc}})$-DP retrieval guarantee against the tenant index.

arXiv Machine Learning
Aug 20

Topology-Aware Differential Privacy in Hierarchical Federated Learning

The paper introduces Fulcrum, a topology‑aware differential privacy scheme for hierarchical federated learning that allocates noise based on the size and exposure of regional aggregation groups. By deriving a closed‑form exposure dispersion metric from region structure and weights, the method optimally balances privacy and utility, achieving up to 14.84% accuracy gains on image tasks and 12.16% on text tasks at ε = 0.99 compared to uniform noise allocation. The approach ensures each participant receives noise commensurate with its actual exposure, eliminating unnecessary privacy overhead.

By Murtaza Rangwala, Richard O. Sinnott, Rajkumar Buyya
arXiv Machine Learning
Sep 10

Characterizing Privacy-Audit Alignment in Behavioral Audit of Machine Unlearning

The paper investigates the privacy risks inherent in auditing machine unlearning (MU) when the audit relies only on querying the model for behavioral signals. It shows that such generic audit schemes inevitably leak information about the retained data set, providing a geometric transfer theorem that bounds the distinguishability of retained set membership based on audit accuracy. The study also analyzes how the unlearned set, target sample, and query protocol influence the privacy‑audit transfer coefficient, with empirical evidence from both convex and non‑convex models supporting the theoretical findings.

By Liou Tang, James Joshi, Ashish Kundu
arXiv AI
Sep 10

PAC-Private Autoregressive Generation: Calibrating Noise to Ensemble Disagreement

The paper introduces PAC‑Private Autoregressive Generation, a method that calibrates noise based on ensemble disagreement across overlapping ‘worlds’ of a private corpus, thereby extending PAC privacy from classification to text generation. By training adapters on a frozen public model and using posterior‑weighted disagreement to add noise only when predictions vary, the approach achieves strong privacy guarantees while preserving most of the fine‑tuning benefit. Experiments on WikiText‑103 with GPT‑2‑small show 74 % of the fine‑tuning gain retained with a per‑token budget of 2⁻³², and membership‑inference success bounded to 51.08 % after one million tokens, outperforming PMixED under matched conditions.

By Mina Mirzadehsarcheshmeh, Amir Keyvan Khandani
arXiv AI
Sep 4

Privacy-Preserving Topology-Guided Safety for LLM-Based Multi-Agent Systems via Federated Graph Learning

The paper introduces FGLGuard, a privacy‑preserving federated graph learning framework that trains a graph attention detector on each operator’s own multi‑agent system (MAS) episode graphs, sharing only model updates. By combining a proximal local objective, domain‑balanced aggregation, threshold calibration, and guarded rewrite mechanisms, FGLGuard adapts to non‑IID data across organizations and outperforms centralized and local‑only baselines on Agent‑SafetyBench, R‑Judge, and AgentDojo. The method achieves significant reductions in attack success rates—up to 43% on AgentDojo—without compromising utility, API cost, or model capability.

By Jinxi Yu, Eric Hanchen Jiang, Levina Li, Dong Liu, Zhi Zhang, Wenxiao Zhao, Yanxuan Yu, Kai-Wei Chang, Ying Nian Wu
arXiv Machine Learning
Sep 23

Optimizing Canaries for Privacy Auditing with Metagradient Descent

The paper investigates black-box privacy auditing for differentially private learning algorithms, focusing on DP‑SGD. It introduces a method that optimizes the auditor’s canary set using metagradient descent, improving empirical lower bounds on privacy parameters compared to prior canary designs. The approach is shown to be DP‑SGD agnostic and efficient, with optimized canaries for small models remaining effective for larger DP‑SGD models.

By Matteo Boglioni, Terrance Liu, Andrew Ilyas, Zhiwei Steven Wu
arXiv AI
Jul 15

Cost-Governed RAG: Unified Per-Tenant Cost Attribution Across Retrieval and Generation in Multi-Tenant LLM Systems

arXiv:2607. 12188v1 Announce Type: new Abstract: Enterprise Retrieval-Augmented Generation (RAG) deployments face a critical governance gap: while LLM generation cost is metered per token, the retrieval layer - vector memory, similarity compute, and embedding API calls - remains an unattributed shared cost, enabling invisible cross-subsidization among tenants.

By Navnit Shukla
arXiv AI
Sep 11

Subgroup Membership Inference Audits of Differentially Private Synthetic Text

The paper introduces a subgroup-targeted membership inference game to audit differentially private synthetic text releases, revealing that existing average-case attacks miss significant leakage to vulnerable subgroups. An extensive audit across 32 proxies, four datasets, three generation methods, and five privacy budgets shows that DP reduces overall leakage but leaves concentrated, uneven residual risk, especially for high-risk records. The study demonstrates that which records leak is determined by the release mechanism rather than the records themselves, challenging record-level risk assessment.

By Yidan Sun, Viktor Schlegel, Srinivasan Nandakumar, Siew Kei Lam, Anil Anthony Bharath