arXiv Machine Learning

Adaptively Robust LLM Monitoring via Activation Watermarking

arXiv:2603. 23171v3 Announce Type: replace-cross Abstract: Providers monitor deployed large language models (LLMs) to detect misuse that they cannot prevent.

arXiv Machine Learning
Jul 8

Multi-Channel Spread-Spectrum Code Watermarking

arXiv:2607. 06009v1 Announce Type: cross Abstract: Attributing code to the large language model that produced it is essential for provenance, licensing, and misuse accountability, yet no deployed watermark meets this need.

By Soohyeon Choi, Debin Gao, Yue Duan
arXiv AI
Jun 4

MaskForge: Structure-Aware Adaptive Attacks for Jailbreaking Diffusion Large Language Models

arXiv:2606. 04027v1 Announce Type: cross Abstract: Diffusion large language models (dLLMs) generate text by iteratively denoising partially masked sequences under bidirectional context, exposing a safety surface distinct from autoregressive LLMs.

By Yingzi Ma, Zhengyue Zhao, Xiaogeng Liu, Minhui Xue, Yue Zhao, Chaowei Xiao
arXiv Computation and Language
Sep 11

SpecGuard: Inference-Time Backdoor Detection For Free

SpecGuard is an inference‑time backdoor detector that leverages speculative decoding—using a small draft model to propose tokens and a target model to verify them—without adding extra model computation. By monitoring the draft‑token acceptance rate, SpecGuard detects when a target model shifts toward attacker‑controlled behavior while the draft model does not, signaling a backdoor trigger. The method reliably identifies a range of backdoor types, including stealthy attacks that bypass input‑level filters, across multiple model families.

By Rui Wen, Ahmed Salem, Andrew Paverd, Mark Russinovich, Zheng Li
arXiv Machine Learning
Jul 30

ToxScreen: Detecting Whether an LLM Has Been Poisoned

arXiv:2607. 26849v1 Announce Type: cross Abstract: As large language models (LLMs) are deployed in high-stakes domains, adversaries may poison training data to implant backdoors: hidden triggers that covertly manipulate model behavior at inference time.

By Anthony Hughes, Nicole Xing, Collin Francel, Andy Kim, Andrew Draganov
arXiv AI
Jun 4

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering

arXiv:2602. 06911v2 Announce Type: replace-cross Abstract: As increasingly capable open-weight large language models (LLMs) are deployed, improving their tamper resistance against unsafe modifications, whether accidental or intentional, becomes critical to minimize risks.

By Saad Hossain, Tom Tseng, Punya Syon Pandey, Samanvay Vajpayee, Matthew Kowal, Nayeema Nonta, Samuel Simko, Stephen Casper, Zhijing Jin, Kellin Pelrine, Sirisha Rambhatla
arXiv AI
Sep 15

Overflip: Repetition-Induced Label Flips in Guardrail Models

Guardrail models, which screen malicious prompts in LLM services, often use lightweight Transformers with short context windows and bucketed positional encodings. The study identifies a new failure mode called Overflip, where repeating a prompt causes the guardrail’s prediction to flip from malicious to benign as the sequence length increases. Experiments on nine popular guardrails show that 5 models exhibit MAL→BEN flips on 100 prompts, with flip rates ranging from 8% to 92% and first flips occurring between 2.6k and 9.4k tokens, highlighting a gradual attention dispersion distinct from traditional attention‑dilution attacks.

By Xu He, Chih-Hsuan Lin, Hung-Mao Chen, Junjie Xiong, Yan Zhai, Kun Sun