The paper introduces the first systematic study of zero‑knowledge (ZK)‑friendly quantization for large language models (LLMs). It defines what makes a quantization scheme suitable for ZK proof generation and evaluates nine models, including Qwen2.5‑14B and Qwen3‑30B‑A3B, across various weight, activation, and nonlinear lookup precisions. Findings reveal that activation precision is more critical than weight precision, nonlinear lookup approximations can dominate utility loss, and that reducing bit‑width or lookup size does not always lead to proportional proving cost savings, highlighting the need for operator‑aware precision selection.
By Taeung Yoon, Yupeng Zhang, Xiaojing Liao
arXiv:2605. 24033v2 Announce Type: replace Abstract: Mechanistic interpretability typically discovers circuits and then argues what they do from examples and ablations.
By Neel Somani
Open-source large language models (LLMs) are increasingly competitive with closed-source models while offering transparency and the ability to run inference without exposing user inputs to a service p...
arXiv:2606. 05433v1 Announce Type: new Abstract: Frontier AI governance frameworks increasingly use cumulative training compute as the primary criterion for designating high-impact models, but enforcement rests on self-reporting because no technical verification primitive for training exists.
By Pierre Peign\'e, Ky Nguyen, Paul Wang
TensorCommitments (TCs) is a lightweight, tensor-native proof‑of‑inference scheme that enables verifiable inference for large language models (LLMs) without requiring the verifier to rerun the model or possess a powerful GPU. By binding each inference to a commitment stored in multivariate Terkle Trees, TCs detect tampering with only a 0.97% overhead for the prover and 0.12% for the verifier on LLaMA2. The approach improves robustness against tailored LLM attacks by up to 48% compared to previous methods that needed a verifier GPU.
By Oguzhan Baser, Elahe Sadeghi, Eric Wang, Nico Vergauwen, Sam Kazemian, Hong Kang, Sandeep P. Chinchali, Sriram Vishwanath
Werracle is a zero‑storage on‑chain AI decision oracle that fits into a single 32‑byte EVM storage slot and uses procedural Mandelbrot dynamics to generate continuous non‑linear decision hyperplanes from a 24‑byte coordinate triplet. Implemented in pure Solidity bytecode with fixed‑point arithmetic, it evaluates a 16‑point Pareto micro‑grid in only 21,438 gas, achieving sub‑millisecond latency on Layer‑2 rollups. The protocol is formally verified with a 1,000‑test deterministic suite and is deployed live on an EVM devnet, demonstrated through a Uniswap v4 dynamic fee governor that adjusts liquidity provider fees in real time.
By Volkan Da\u{g}l{\i}, Zerrin Da\u{g}l{\i}, Da\u{g}han Da\u{g}l{\i}