arXiv Machine Learning

Detecting Hidden ML Training With Zero-Overhead Telemetry

arXiv:2606. 19262v1 Announce Type: new Abstract: Hardware-enabled monitoring of GPU workloads underpins many proposals for AI compute governance, but if developers can defeat monitoring mechanisms, such schemes are unworkable.

arXiv AI
Sep 2

Workload Identification with Physical Side Channels for AI Governance

The paper demonstrates that an external observer can identify the type of workload running on an NVIDIA H200 GPU by analyzing its power draw, distinguishing training, inference, and non‑AI tasks with high accuracy. Using 930 recorded traces, the authors achieve 97% accuracy and a macro‑averaged F1 score of 0.955 on unseen model families. They also test four evasion strategies to disguise training as inference, showing that a hardened detector can catch most attacks, though one strategy (LoRA) remains partially detectable.

By Simone Gargiulo, Gabriel Kulp
arXiv Machine Learning
Jun 2

Bit-Exact AI Inference Verification Without Performance Tradeoffs

arXiv:2606. 00279v1 Announce Type: cross Abstract: Verifying claims about AI workloads is a pre- requisite for credible AI governance of covert adversaries (who comply with monitoring only when detection likelihood is high), yet the ap- parent non-determinism of GPU floating-point arithmetic forces auditors to accept approximate output matches.

By Naci Cankaya
arXiv Machine Learning
Sep 16

OPEN-1B: A Fully Auditable Training Run

The paper introduces Open-1B, a language model trained under a new fully auditable regime that ensures every training operation is reproducible on heterogeneous commodity hardware with bitwise certainty. By enforcing a fixed order on sources of nondeterminism—GPU reductions, data batch ordering, and inter/intra-node communication—the authors enable auditors to replay and verify individual training steps on a single machine. The release includes the full pretraining dataset, all intermediate checkpoints, the training codebase, and an audit harness for step-by-step verification.

By John Donaghy, Brian Wilcox, O\u{g}uzhan Ersoy, Shikhar Rastogi, Adam St Arnaud, Alexey Titov, Jordan Greenberg, Ben Fielding, Harry Grieve
arXiv Machine Learning
Aug 31

Not to Break, but to Attest: Adversarial Probes for Privacy-Preserving LLM Verification

The paper introduces a privacy‑preserving zk‑SNARK audit framework that uses adversarial‑style probes to detect logit drift between an approved large language model and a modified deployment. It offers three probe families—token‑based (black‑box), embedding‑based (gray‑box), and stress probes (partial white‑box)—allowing users to balance sensitivity, access, and cost. Experiments across LLM architectures and GPU platforms show token‑based probes achieve the highest mean sensitivity while remaining practical in a black‑box setting, with Groth16 proving times scaling modestly from 1.02 to 1.78 seconds and constant proof size.

By Cameron Wilding, Mina Shaker, Fatemeh Ganji
arXiv AI
Sep 21

TERMon: Detecting Persistent Behavioral Threats in Edge AI via Hardware-Native Ternary Runtime Monitor

TERMon is a lightweight hardware runtime monitor designed for edge AI accelerators in safety‑critical environments. It detects persistent behavioral threats—such as model corruption, distribution shift, and adversarial inputs—by observing inference behavior through hardware‑efficient ternary patterns matched against a thermometer‑encoded fingerprint. Implemented on a PYNQ‑Z2 FPGA, TERMon operates with a two‑cycle decision latency and requires no on‑chip block RAM or DSPs.

By Arish Sateesan, Edlira Dushku
arXiv Machine Learning
Aug 24

RouteScan: A Non-Intrusive Approach to Auditing MoE LLMs Safety via Expert Routing Telemetry

RouteScan is a non‑intrusive auditing framework that detects harmful behavior in Mixture‑of‑Experts (MoE) large language models by analyzing expert‑routing telemetry captured from GPU execution. It uses the number of active GPU threads during the prefilling phase as a micro‑architectural fingerprint to isolate cross‑domain risk indicators and precisely identify malicious prompts. Evaluations on four open‑source MoE LLMs show strong generalization with AUROC > 0.91 on unseen harmful domains, while privacy tests indicate that full prompts cannot be reliably recovered from aggregated telemetry.

By Bo Lv, Zhiheng Xu, KeDong Xiu, Ruyi Ding, Tianhang Zheng, Zhibo Wang, Kui Ren
arXiv Machine Learning
4d ago

OVIG: Optimistic Verification of AI Training Integrity via Gradient Signals

OVIG is an optimistic verification framework that audits AI training by replaying the process and comparing gradient differences against an empirically calibrated boundary. It treats any gradient difference exceeding this boundary as a malicious deviation. By partitioning training into stride‑s intervals and storing evidence only at interval endpoints, OVIG dramatically reduces off‑chain storage and transmission costs while maintaining zero attack success rate across language, vision, and diffusion workloads.

By Hongxu Su, Jianzhu Yao, Huan Zhang, Xuechao Wang, Pramod Viswanath
arXiv AI
Sep 15

SpliTEE: Improving LLM Inference on Trusted Hardware with Differentially Private GPU Outsourcing

SpliTEE extends the split‑inference architecture of Slalom to large language models by protecting intermediate GPU computations with differential privacy rather than encryption. The authors show that masking intermediate representations is essential, as a prompt‑reconstruction attack can recover prompts with about 80% accuracy. Their global sensitivity analysis bounds the noise needed, and they demonstrate that SpliTEE on Intel TDX achieves near‑double the speed of fully CPU‑based inference and outperforms encryption‑based Slalom while maintaining higher accuracy.

By Shashie Dilhara Batan Arachchige, Robin Carpentier, Hassan Jameel Asghar, Dali Kaafar
arXiv Machine Learning
Aug 24

Faults That Fortify: CNN Adversarial Robustness via GPU Undervolting

The paper demonstrates that undervolting GPUs during CNN training introduces stochastic faults that act as implicit regularization, improving adversarial robustness while reducing power consumption. Experiments on LeNet, VGG-6, and MobileNetV3 trained on MNIST and CIFAR-10 show that undervolted models consistently outperform nominal-voltage models in both standard and adversarial training regimes. The approach offers a hardware-level defense that requires no algorithmic changes and yields significant energy savings due to the quadratic relationship between dynamic power and supply voltage.

By Behnam Omidi, Ahmad Tahmasivand, Husam Alsyouri, Saba Al-Sayouri, Chongzhou Fang, Ihsen Alouani, Khaled N. Khasawneh