arXiv AI

NanoZK: Privacy-Preserving Verifiable Inference for Large Language Models via Layerwise Zero-Knowledge Proofs

arXiv:2603. 18046v2 Announce Type: replace-cross Abstract: We present NanoZK, a zero-knowledge proof system for verifiable LLM inference: clients and third-party auditors check that a provider executed the advertised model on a committed input without learning weights or activations.

arXiv AI
4d ago

Bits Under ZK-LLM: Evaluating Zero-Knowledge-Friendly Quantization for Verifiable Private LLM Inference

The paper introduces the first systematic study of zero‑knowledge (ZK)‑friendly quantization for large language models (LLMs). It defines what makes a quantization scheme suitable for ZK proof generation and evaluates nine models, including Qwen2.5‑14B and Qwen3‑30B‑A3B, across various weight, activation, and nonlinear lookup precisions. Findings reveal that activation precision is more critical than weight precision, nonlinear lookup approximations can dominate utility loss, and that reducing bit‑width or lookup size does not always lead to proportional proving cost savings, highlighting the need for operator‑aware precision selection.

By Taeung Yoon, Yupeng Zhang, Xiaojing Liao
arXiv AI
Jun 6

Zero knowledge verification for frontier AI training is possible

arXiv:2606. 05433v1 Announce Type: new Abstract: Frontier AI governance frameworks increasingly use cumulative training compute as the primary criterion for designating high-impact models, but enforcement rests on self-reporting because no technical verification primitive for training exists.

By Pierre Peign\'e, Ky Nguyen, Paul Wang
arXiv AI
2d ago

TensorCommitments: A Lightweight Verifiable Inference for Language Models

TensorCommitments (TCs) is a lightweight, tensor-native proof‑of‑inference scheme that enables verifiable inference for large language models (LLMs) without requiring the verifier to rerun the model or possess a powerful GPU. By binding each inference to a commitment stored in multivariate Terkle Trees, TCs detect tampering with only a 0.97% overhead for the prover and 0.12% for the verifier on LLaMA2. The approach improves robustness against tailored LLM attacks by up to 48% compared to previous methods that needed a verifier GPU.

By Oguzhan Baser, Elahe Sadeghi, Eric Wang, Nico Vergauwen, Sam Kazemian, Hong Kang, Sandeep P. Chinchali, Sriram Vishwanath
arXiv AI
6d ago

Werracle: Sub-Cent Intra-Block AI Reflex Oracles and Flash-Loan Circuit Breakers for EVM Smart Contracts

Werracle is a zero‑storage on‑chain AI decision oracle that fits into a single 32‑byte EVM storage slot and uses procedural Mandelbrot dynamics to generate continuous non‑linear decision hyperplanes from a 24‑byte coordinate triplet. Implemented in pure Solidity bytecode with fixed‑point arithmetic, it evaluates a 16‑point Pareto micro‑grid in only 21,438 gas, achieving sub‑millisecond latency on Layer‑2 rollups. The protocol is formally verified with a 1,000‑test deterministic suite and is deployed live on an EVM devnet, demonstrated through a Uniswap v4 dynamic fee governor that adjusts liquidity provider fees in real time.

By Volkan Da\u{g}l{\i}, Zerrin Da\u{g}l{\i}, Da\u{g}han Da\u{g}l{\i}
arXiv AI
Sep 15

SpliTEE: Improving LLM Inference on Trusted Hardware with Differentially Private GPU Outsourcing

SpliTEE extends the split‑inference architecture of Slalom to large language models by protecting intermediate GPU computations with differential privacy rather than encryption. The authors show that masking intermediate representations is essential, as a prompt‑reconstruction attack can recover prompts with about 80% accuracy. Their global sensitivity analysis bounds the noise needed, and they demonstrate that SpliTEE on Intel TDX achieves near‑double the speed of fully CPU‑based inference and outperforms encryption‑based Slalom while maintaining higher accuracy.

By Shashie Dilhara Batan Arachchige, Robin Carpentier, Hassan Jameel Asghar, Dali Kaafar
arXiv Machine Learning
Aug 19

Certified but Private: Scalable Zero-Knowledge Proofs for Neural Network Guarantees

PANDA is a scalable system that uses zero‑knowledge proofs to certify the robustness and fairness of neural networks without revealing their private parameters. Built on the CROWN robustness framework, PANDA introduces a novel algorithm for proving linear relaxation bounds on non‑linear activation layers, producing lightweight proofs. The system can generate proofs for networks with over 2.9 million parameters in just five minutes and verify them in ten seconds, scaling polynomially with network size and enabling verification of models four orders of magnitude larger than prior ZKP‑based approaches.

By Youwei Zhong, Ben Merbaum, Timos Antonopoulos, Ning Luo, Charalampos Papamanthou, Katerina Sotiraki, Ruzica Piskac