The paper introduces a privacy‑preserving zk‑SNARK audit framework that uses adversarial‑style probes to detect logit drift between an approved large language model and a modified deployment. It offers three probe families—token‑based (black‑box), embedding‑based (gray‑box), and stress probes (partial white‑box)—allowing users to balance sensitivity, access, and cost. Experiments across LLM architectures and GPU platforms show token‑based probes achieve the highest mean sensitivity while remaining practical in a black‑box setting, with Groth16 proving times scaling modestly from 1.02 to 1.78 seconds and constant proof size.
By Cameron Wilding, Mina Shaker, Fatemeh Ganji
arXiv:2606. 11998v1 Announce Type: new Abstract: Trusted monitoring is a cornerstone of AI control.
By Frank Xiao, Mary Phuong
TensorCommitments (TCs) is a lightweight, tensor-native proof‑of‑inference scheme that enables verifiable inference for large language models (LLMs) without requiring the verifier to rerun the model or possess a powerful GPU. By binding each inference to a commitment stored in multivariate Terkle Trees, TCs detect tampering with only a 0.97% overhead for the prover and 0.12% for the verifier on LLaMA2. The approach improves robustness against tailored LLM attacks by up to 48% compared to previous methods that needed a verifier GPU.
By Oguzhan Baser, Elahe Sadeghi, Eric Wang, Nico Vergauwen, Sam Kazemian, Hong Kang, Sandeep P. Chinchali, Sriram Vishwanath
arXiv:2608.21803v1 Announce Type: cross
Abstract: As machine learning (ML) models are increasingly deployed in high-stakes environments, explainable AI (XAI) methods like SHAP and LIME have become es...
By Maraz Mia, Shovan Roy, Mir Mehedi A. Pritom, Maanak Gupta
arXiv:2604. 04738v2 Announce Type: replace-cross Abstract: Fine-tuning is the dominant paradigm for adapting large machine learning models, yet current deployment pipelines provide no way to verify how a released model was updated.
By Zhenhang Shang, Yingzhe Yu, Kani Chen
The paper introduces Open-1B, a language model trained under a new fully auditable regime that ensures every training operation is reproducible on heterogeneous commodity hardware with bitwise certainty. By enforcing a fixed order on sources of nondeterminism—GPU reductions, data batch ordering, and inter/intra-node communication—the authors enable auditors to replay and verify individual training steps on a single machine. The release includes the full pretraining dataset, all intermediate checkpoints, the training codebase, and an audit harness for step-by-step verification.
By John Donaghy, Brian Wilcox, O\u{g}uzhan Ersoy, Shikhar Rastogi, Adam St Arnaud, Alexey Titov, Jordan Greenberg, Ben Fielding, Harry Grieve
arXiv:2603. 05786v2 Announce Type: replace-cross Abstract: As AI agents become widely deployed as online services, users often rely on an agent developer's claim about how safety is enforced, which introduces a threat where safety measures are falsely advertised.
By Xisen Jin, Michael Duan, Qin Lin, Aaron Chan, Zhenglun Chen, Junyi Du, Xiang Ren
arXiv:2606. 16358v1 Announce Type: cross Abstract: Agents increasingly access large language models (LLMs) through API routers.
By Sipeng Xie, Qianhong Wu, Hengrun Lu, Ziliang Sun, Qi Wu, Bo Qin, Qin Wang
The paper investigates how inference optimization for large language models can introduce numerical inconsistencies that trigger hidden backdoors. It introduces two types of optimization‑triggered backdoors: the Input‑Specific Optimization Backdoor (ISOB) and the Universal Optimization Backdoor (UOB), the latter enabling a model to remain benign under normal execution but activate a backdoor when optimization is applied. Experiments on seven open‑source LLMs, across multiple tasks and optimization backends, show UOB can achieve up to 100% attack success while maintaining clean accuracy, and the authors propose three defenses that reduce the attack success rate to 0.02.
By Yifei Wang, Yida Yang, Tianlin Li, Xiaohan Zhang, Xiaoyu Zhang, Li Pan
arXiv:2604. 01039v2 Announce Type: replace-cross Abstract: System Instructions in Large Language Models (LLMs) are commonly used to enforce safety policies, define agent behavior, and protect sensitive operational context in agentic AI applications.
By Anubhab Sahu, Diptisha Samanta, Reza Soosahabi
arXiv:2510. 01529v3 Announce Type: replace Abstract: Ball et al.
By Jaiden Fairoze, Sanjam Garg, Keewoo Lee, Mingyuan Wang
OVIG is an optimistic verification framework that audits AI training by replaying the process and comparing gradient differences against an empirically calibrated boundary. It treats any gradient difference exceeding this boundary as a malicious deviation. By partitioning training into stride‑s intervals and storing evidence only at interval endpoints, OVIG dramatically reduces off‑chain storage and transmission costs while maintaining zero attack success rate across language, vision, and diffusion workloads.
By Hongxu Su, Jianzhu Yao, Huan Zhang, Xuechao Wang, Pramod Viswanath