arXiv AI

SafeSpec: Fast and Safe LLM via Dynamic Reflective Sampling

arXiv:2606. 19755v1 Announce Type: cross Abstract: Speculative inference accelerates large language model (LLM) decoding but provides no inherent safety guarantees.

Hugging Face Trending Papers
Jul 2

Safety Targeted Embedding Exploit via Refinement

Safety training for large language models (LLMs) is conducted predominantly in English, leaving uncertain how well safety mechanisms generalize to low-resource languages and mixed-language code-switching. We show that this creates an epistemic gap in which models confidently generate harmful responses for inputs that fall outside the distribution of their safety training.

arXiv Machine Learning
Sep 11

Understanding In-Context Multimodal Jailbreaks via Posterior Reweighting

The paper introduces a posterior reweighting framework to explain and counter in-context learning jailbreaks in multimodal large language models. It models the model as switching between safe and harmful behavioral modes, interpreting prompt demonstrations as evidence that shifts the posterior. Using this view, the authors derive scaling laws for jailbreak effectiveness and propose a defense that injects benign counter‑evidence to suppress harmful drift while maintaining utility.

By Xu Zhang, Dev Mistry, Xiang Xu, Ren Wang
arXiv Machine Learning
Jul 27

Adversarial Prompts for Acceptance Collapse in Speculative Decoding

arXiv:2607. 21804v1 Announce Type: cross Abstract: Lossless acceleration schemes, such as speculative decoding, promise significant inference speedups by relying on dynamic token-level alignment between a draft and a target model.

By Run Wang, Chaoyi Zhou, Xi Liu, Yi Zhu, Amir Salarpour, Pedram MohajerAnsari, Zhi-Qi Cheng, Feng Luo, Siyu Huang, Mert D. Pes\'e
arXiv Computation and Language
Sep 11

SpecGuard: Inference-Time Backdoor Detection For Free

SpecGuard is an inference‑time backdoor detector that leverages speculative decoding—using a small draft model to propose tokens and a target model to verify them—without adding extra model computation. By monitoring the draft‑token acceptance rate, SpecGuard detects when a target model shifts toward attacker‑controlled behavior while the draft model does not, signaling a backdoor trigger. The method reliably identifies a range of backdoor types, including stealthy attacks that bypass input‑level filters, across multiple model families.

By Rui Wen, Ahmed Salem, Andrew Paverd, Mark Russinovich, Zheng Li
arXiv Computation and Language
Sep 7

Revisiting Lossy Verification in Speculative Decoding: Mechanisms, Trade-offs, and Failure Modes

The paper examines lossy verification techniques used in speculative decoding for large language models, showing that many methods can be grouped into truncation-based and collaborative verification categories. It analyzes how these approaches alter the decoding distribution, revealing that truncation-based methods can significantly degrade performance due to distributional distortion, while collaborative methods depend more on overshoot suppression and supervision quality than on simple interpolation between draft and target models. A diagnostic evaluation framework is introduced to assess these failure modes across curated benchmarks.

By Tianyu Wang, Yuxuan Zhou, Heng Li, Wenbin Wang, Zikai Xiao, Chunrui Zheng, Junyuan Shang