arXiv AI

Black-Box Inference of LLM Architectural Properties with Restrictive API Access

arXiv:2607. 01313v1 Announce Type: cross Abstract: In practice, most commercial LLM providers do not publicly release details of underlying LLM architectures.

arXiv Machine Learning
Aug 31

Not to Break, but to Attest: Adversarial Probes for Privacy-Preserving LLM Verification

The paper introduces a privacy‑preserving zk‑SNARK audit framework that uses adversarial‑style probes to detect logit drift between an approved large language model and a modified deployment. It offers three probe families—token‑based (black‑box), embedding‑based (gray‑box), and stress probes (partial white‑box)—allowing users to balance sensitivity, access, and cost. Experiments across LLM architectures and GPU platforms show token‑based probes achieve the highest mean sensitivity while remaining practical in a black‑box setting, with Groth16 proving times scaling modestly from 1.02 to 1.78 seconds and constant proof size.

By Cameron Wilding, Mina Shaker, Fatemeh Ganji
arXiv Computation and Language
Sep 1

Token Counts Are Not Model Lineage: A Frozen-Threshold Holdout Study of Black-Box LLM API Fingerprinting

The study evaluates whether prompt‑token counts can reliably identify the lineage of large language models served via APIs. Using a frozen‑threshold approach on 24 labeled endpoint pairs, the authors find that token‑count consistency perfectly separates development pairs but only half of the holdout pairs meet the strict repeatability criteria, yielding moderate accuracy and perfect specificity. The results confirm token‑count consistency as a fingerprint of shared tokenization stacks but reject it as a standalone test for model‑family attribution.

By Bo Chen
arXiv AI
Sep 10

Detokenization Leaks: Reconstructing Local LLM Outputs From Cache Traces

The paper introduces an attack that reconstructs text generated by locally hosted large language models by monitoring CPU cache activity during detokenization. It uses Flush+Reload on shared tokenizer code to time decoding, then Prime+Probe to capture token‑dependent cache traces, followed by a clustering‑and‑language‑model pipeline to recover the output text. The method is evaluated across various datasets, hardware, inference frameworks, and model families, successfully retrieving semantically accurate outputs from real‑world local LLM deployments, including agentic systems.

By Roy Weiss, Benyamin Konstantinov, Eitam Sheetrit, Tomer Simon, Yisroel Mirsky
arXiv Machine Learning
Jul 14

One Token Is Enough: Fingerprinting and Verifying Large Language Models from Single-Token Output Distributions

arXiv:2607. 10252v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly consumed through opaque serving chains - API aggregators, resellers, and inference providers - in which the client has no technical means to confirm that the model answering is the model advertised, and recent audits show that a substantial fraction of commercial endpoints deviate from the vendor's reference weights.

By Tomas Bruckner
arXiv Computation and Language
Sep 11

SpecGuard: Inference-Time Backdoor Detection For Free

SpecGuard is an inference‑time backdoor detector that leverages speculative decoding—using a small draft model to propose tokens and a target model to verify them—without adding extra model computation. By monitoring the draft‑token acceptance rate, SpecGuard detects when a target model shifts toward attacker‑controlled behavior while the draft model does not, signaling a backdoor trigger. The method reliably identifies a range of backdoor types, including stealthy attacks that bypass input‑level filters, across multiple model families.

By Rui Wen, Ahmed Salem, Andrew Paverd, Mark Russinovich, Zheng Li
arXiv AI
Jun 2

AgentRedBench: Dynamic Redteaming and Integration-Aware Defense for LLM Agents over SaaS Integrations

arXiv:2606. 02240v1 Announce Type: cross Abstract: Indirect prompt injection in tool-use agents is a concrete production threat: LLM agents read from integrations (third-party services such as Gmail, Salesforce, or Jira accessed through tool calls) whose response content the user neither writes nor controls.

By Hiskias Dingeto, Will Leeney