arXiv Machine Learning

Rendering on Real Silicon: GPU Render-Timing as a Passive, AI-Resistant CAPTCHA Signal

arXiv:2607. 23389v1 Announce Type: cross Abstract: Conventional CAPTCHAs pose puzzles that modern AI systems increasingly solve, while behavioral and cryptographic-attestation defenses carry privacy or enrollment costs.

arXiv Machine Learning
Jun 2

Bit-Exact AI Inference Verification Without Performance Tradeoffs

arXiv:2606. 00279v1 Announce Type: cross Abstract: Verifying claims about AI workloads is a pre- requisite for credible AI governance of covert adversaries (who comply with monitoring only when detection likelihood is high), yet the ap- parent non-determinism of GPU floating-point arithmetic forces auditors to accept approximate output matches.

By Naci Cankaya
arXiv AI
Aug 25

TEE-X: TEE-aware Acceleration Framework for Large Vision Models at the Edge

TEE-X is a TEE‑aware acceleration framework designed to run large vision models, such as Vision Transformers, entirely within Trusted Execution Environments. It introduces a sensitivity‑aware modularization technique and vectorization to overcome memory constraints and latency challenges on edge devices. The framework is validated on OP‑TEE for Arm TrustZone and optimized for the NVIDIA Jetson AGX Xavier, achieving GPU‑level inference latency with minimal accuracy‑latency trade‑offs.

By Kurt M Wilson, Mohaiminul Al Nahian, Abeer Matar A. Almalky, Sadat Shahriyar, Souvik Kundu, Zhishan Guo, Abdullah Al Arafat, Adnan Siraj Rakin
Hugging Face Trending Papers
5d ago

AgentPerfBench: A Benchmarking and Evaluation Suite for Inference Performance of Agentic LLMs

AgentPerfBench is a new benchmarking suite designed to evaluate the inference performance of agentic large language models (LLMs) that handle multi‑turn, tool‑using, and context‑expanding tasks. It builds on real traces from agentic benchmarks such as SWE‑Bench and TerminalBench, and generates synthetic profiles that reflect realistic input/output lengths and turn counts. The suite also provides kernel‑level Nsight Compute traces and a multi‑dimensional roofline model to identify hardware bottlenecks and quantify the gap between traditional chat benchmarks and agentic workloads.

arXiv Machine Learning
Aug 24

RouteScan: A Non-Intrusive Approach to Auditing MoE LLMs Safety via Expert Routing Telemetry

RouteScan is a non‑intrusive auditing framework that detects harmful behavior in Mixture‑of‑Experts (MoE) large language models by analyzing expert‑routing telemetry captured from GPU execution. It uses the number of active GPU threads during the prefilling phase as a micro‑architectural fingerprint to isolate cross‑domain risk indicators and precisely identify malicious prompts. Evaluations on four open‑source MoE LLMs show strong generalization with AUROC > 0.91 on unseen harmful domains, while privacy tests indicate that full prompts cannot be reliably recovered from aggregated telemetry.

By Bo Lv, Zhiheng Xu, KeDong Xiu, Ruyi Ding, Tianhang Zheng, Zhibo Wang, Kui Ren
arXiv AI
Sep 21

TERMon: Detecting Persistent Behavioral Threats in Edge AI via Hardware-Native Ternary Runtime Monitor

TERMon is a lightweight hardware runtime monitor designed for edge AI accelerators in safety‑critical environments. It detects persistent behavioral threats—such as model corruption, distribution shift, and adversarial inputs—by observing inference behavior through hardware‑efficient ternary patterns matched against a thermometer‑encoded fingerprint. Implemented on a PYNQ‑Z2 FPGA, TERMon operates with a two‑cycle decision latency and requires no on‑chip block RAM or DSPs.

By Arish Sateesan, Edlira Dushku
arXiv AI
Sep 2

Workload Identification with Physical Side Channels for AI Governance

The paper demonstrates that an external observer can identify the type of workload running on an NVIDIA H200 GPU by analyzing its power draw, distinguishing training, inference, and non‑AI tasks with high accuracy. Using 930 recorded traces, the authors achieve 97% accuracy and a macro‑averaged F1 score of 0.955 on unseen model families. They also test four evasion strategies to disguise training as inference, showing that a hardened detector can catch most attacks, though one strategy (LoRA) remains partially detectable.

By Simone Gargiulo, Gabriel Kulp