Bits Under ZK-LLM: Evaluating Zero-Knowledge-Friendly Quantization for Verifiable Private LLM Inference
Read the original on arXiv AI →The paper introduces the first systematic study of zero‑knowledge (ZK)‑friendly quantization for large language models (LLMs). It defines what makes a quantization scheme suitable for ZK proof generation and evaluates nine models, including Qwen2.5‑14B and Qwen3‑30B‑A3B, across various weight, activation, and nonlinear lookup precisions. Findings reveal that activation precision is more critical than weight precision, nonlinear lookup approximations can dominate utility loss, and that reducing bit‑width or lookup size does not always lead to proportional proving cost savings, highlighting the need for operator‑aware precision selection.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.