arXiv AI

Policy-Masked Private Experts: Auditable and Reversible Capability Access Control in Sparse MoE Models

arXiv:2608. 06690v1 Announce Type: cross Abstract: Most language-model access controls regulate behavior while leaving the same computation available to every request.

arXiv Machine Learning
Sep 10

RAPTOR: Role-Aware Private Training for Mixture-of-Experts

arXiv:2609.05770v1 Announce Type: new Abstract: Differentially private (DP) fine-tuning methods treat sparse Mixture-of-Experts (MoE) models as a single dense block, ignoring that shared layers see a...

By Duc Dm, Khai Le-Duc, Nguyen Do, Minh Son Hoang, Florent Draye, Thai Hoang, Hoang Phuong Dam, Jiarui Liu, Chris Ngo, Terry Jingchen Zhang, Anh Le Duc Tran, Nhat Do Minh, Minh Ngoc Le, My T. Thai, Ran Xu, Silvio Savarese, Mona Diab, Bernhard Sch\"olkopf, Zhijing Jin, Huy L. Nguyen, Daeyoung Kim
arXiv Machine Learning
Sep 21

ServeGuard: Verifiable, Bounded-Residual Confinement of Operator-Invisible Channels Without Revealing the Certified Read Factor

ServeGuard is a supply‑chain primitive that allows a publisher to ship a proof‑carrying adapter for an open‑weight language model, proving in zero‑knowledge that the adapter contains no hidden backdoor channel in the monitor’s blind subspace. The proof is inexpensive because it relies on a deterministic function of the public base model, and the served residual is the model’s own public floor. The system lets consumers or regulators verify the absence of this class of hidden channels without revealing the certified read factor or trusting the publisher.

By Dominik Dahlem, Rui Vieira
arXiv Machine Learning
Sep 7

Privacy Failure in Split-LLM Training, The Returned Gradient Nullifies the Decoys

The paper reports a privacy breach in a two-node split‑LLM training system where the returned gradient reveals which data rows were real, despite the system passing standard privacy checks. By exploiting the fact that decoy rows produce zero gradients, an attacker can identify real rows with 100% accuracy across multiple runs. The authors demonstrate that adding gradient clipping and noise can mitigate the leak, but the system remains vulnerable to several untested attack vectors.

By Georgios Politis, Evangelos Pappas