arXiv AI By Oguzhan Baser, Elahe Sadeghi, Eric Wang, Nico Vergauwen, Sam Kazemian, Hong Kang, Sandeep P. Chinchali, Sriram Vishwanath

TensorCommitments: A Lightweight Verifiable Inference for Language Models

Read the original on arXiv AI →

TensorCommitments (TCs) is a lightweight, tensor-native proof‑of‑inference scheme that enables verifiable inference for large language models (LLMs) without requiring the verifier to rerun the model or possess a powerful GPU. By binding each inference to a commitment stored in multivariate Terkle Trees, TCs detect tampering with only a 0.97% overhead for the prover and 0.12% for the verifier on LLaMA2. The approach improves robustness against tailored LLM attacks by up to 48% compared to previous methods that needed a verifier GPU.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 15

SpliTEE: Improving LLM Inference on Trusted Hardware with Differentially Private GPU Outsourcing

SpliTEE extends the split‑inference architecture of Slalom to large language models by protecting intermediate GPU computations with differential privacy rather than encryption. The authors show that masking intermediate representations is essential, as a prompt‑reconstruction attack can recover prompts with about 80% accuracy. Their global sensitivity analysis bounds the noise needed, and they demonstrate that SpliTEE on Intel TDX achieves near‑double the speed of fully CPU‑based inference and outperforms encryption‑based Slalom while maintaining higher accuracy.

By Shashie Dilhara Batan Arachchige, Robin Carpentier, Hassan Jameel Asghar, Dali Kaafar
arXiv Machine Learning
Jun 2

Bit-Exact AI Inference Verification Without Performance Tradeoffs

arXiv:2606. 00279v1 Announce Type: cross Abstract: Verifying claims about AI workloads is a pre- requisite for credible AI governance of covert adversaries (who comply with monitoring only when detection likelihood is high), yet the ap- parent non-determinism of GPU floating-point arithmetic forces auditors to accept approximate output matches.

By Naci Cankaya