arXiv:2608.31142v1 Announce Type: cross
Abstract: The 2025--2026 AI market has seen a wave of stealth releases: frontier models launched anonymously on developer platforms under codenames. For their...
By Yisen Xi
arXiv:2606. 00152v1 Announce Type: cross Abstract: LLM-based agents are rapidly advancing, autonomously invoking external tools to complete multi-step tasks for users.
By Mingxuan Zhang, Jiahui Han, Dadi Guo, Songze Li, Guanchu Wang, Na Zou, Dongrui Liu, Xia Hu
arXiv:2609.14003v1 Announce Type: cross
Abstract: Personal AI agents built on large language models (LLMs) are increasingly given access to a user's private data and communications in order to provid...
By Minsun Shim, Ramisha Raida Karim, Ruthwik Jakkula, Kaiwen Zhou, Xin Liu, Xin Eric Wang, Zhou Li
arXiv:2606. 16952v2 Announce Type: replace-cross Abstract: The rapid adoption of generative AI and Large Language Models (LLMs) has spurred interest in synthetic data as a privacy-preserving alternative to sensitive real-world datasets.
By Kareem Amin, Rudrajit Das, Alessandro Epasto, Adel Javanmard, Dennis Kraft, M\'onica Ribero, Sergei Vassilvitskii
arXiv:2606. 16952v1 Announce Type: cross Abstract: The rapid adoption of generative AI and Large Language Models (LLMs) has spurred interest in synthetic data as a privacy-preserving alternative to sensitive real-world datasets.
By Kareem Amin, Rudrajit Das, Alessandro Epasto, Adel Javanmard, Dennis Kraft, M\'onica Ribero, Sergei Vassilvitskii
Tokenized Key-Gated Adapter Routing (Locket) is a framework that embeds fine‑grained, policy‑driven access control into large language models by training lightweight LoRA adapters for different privacy policies. A gating module associates a learned keyed entry token with a specific adapter, allowing authorized tokens to unlock private knowledge while invalid or missing tokens trigger privacy‑preserving adapters that redact or sanitize sensitive content. Experiments on datasets such as Enron, ECHR, and Yelp with models like Qwen3, Llama‑3.2, and Gemma‑2‑2B show that Locket maintains perplexity comparable to fine‑tuning when the correct token is provided, and significantly reduces PII leakage when the token is absent or invalid, without sacrificing utility.
By Mohamed Shaaban, Mohamed Elmahallawy
The paper surveys 25 studies that use explainable AI to compromise machine learning models, covering attacks such as model extraction, membership inference, and model inversion. It distinguishes between how explanations are obtained—through target releases, attacker-derived methods, secondary disclosure, privileged access, or global artifacts—and shows that explanations can lower extraction costs and reveal membership signals via statistics, recourse distance, and robustness. The authors compare threat models, signals, and defenses, concluding that no single explanation type is always unsafe and that protection must be tailored to the specific acquisition path and target asset.
By Abdullah Caglar Oksuz, Anisa Halimi, Erman Ayday
arXiv:2601. 17360v2 Announce Type: replace-cross Abstract: An adversary observing a model's released prediction can infer sensitive attributes of the queried input, or even reconstruct representatives of the model's training data.
By Jiankai Jin, Xiangzheng Zhang, Zhao Liu, Wenzhuo Xu, Dongdong Yang, Deyue Zhang, Quanchen Zou
arXiv:2606. 05433v1 Announce Type: new Abstract: Frontier AI governance frameworks increasingly use cumulative training compute as the primary criterion for designating high-impact models, but enforcement rests on self-reporting because no technical verification primitive for training exists.
By Pierre Peign\'e, Ky Nguyen, Paul Wang
arXiv:2609.38934v1 Announce Type: cross
Abstract: Differentially private (DP) text generation can protect individual records, but privacy alone does not specify what evidence a released statement car...
By Tsubasa Takahashi, Takumi Hiraoka
arXiv:2607. 25364v1 Announce Type: new Abstract: Tool-using agents expose structured calls but commonly attach free-form rationales.
By Genliang Zhu (Accentrust, Georgia Institute of Technology), Chu Wang (Accentrust, University of Illinois Urbana-Champaign)
arXiv:2603. 18046v2 Announce Type: replace-cross Abstract: We present NanoZK, a zero-knowledge proof system for verifiable LLM inference: clients and third-party auditors check that a provider executed the advertised model on a committed input without learning weights or activations.
By Zhaohui Wang