arXiv:2606. 20502v1 Announce Type: cross Abstract: Whether LLMs scoring well on vulnerability benchmarks genuinely reason about security or merely pattern-match on contaminated data remains unresolved.
By Arastoo Zibaeirad, Marco Vieira
The paper introduces TRIM, a black‑box defense for backdoor attacks in computer vision models. TRIM identifies and removes malicious trigger regions at inference time using region‑based segmentation, adaptive trigger discovery via inpainting and diffusion, and selective purification, without needing model internals, training data, or clean samples. Experiments on various datasets and trigger types show TRIM reduces attack success rates to as low as 1.16% while maintaining high clean accuracy.
By Ahmed Abdelnaby, Mohamed Elmahallawy
The paper investigates how inference optimization for large language models can introduce numerical inconsistencies that trigger hidden backdoors. It introduces two types of optimization‑triggered backdoors: the Input‑Specific Optimization Backdoor (ISOB) and the Universal Optimization Backdoor (UOB), the latter enabling a model to remain benign under normal execution but activate a backdoor when optimization is applied. Experiments on seven open‑source LLMs, across multiple tasks and optimization backends, show UOB can achieve up to 100% attack success while maintaining clean accuracy, and the authors propose three defenses that reduce the attack success rate to 0.02.
By Yifei Wang, Yida Yang, Tianlin Li, Xiaohan Zhang, Xiaoyu Zhang, Li Pan
arXiv:2606. 01695v1 Announce Type: new Abstract: Adversaries can implant latent harmful behavior by poisoning as few as 1% of fine-tuning examples.
By Swapnil Parekh
arXiv:2602. 14161v2 Announce Type: replace Abstract: Detecting prompt injection, jailbreak attacks, and harmful requests is critical for deploying LLM-based agents safely, yet current evaluation practices in this literature overestimate generalization.
By Max Fomin
SpecGuard is an inference‑time backdoor detector that leverages speculative decoding—using a small draft model to propose tokens and a target model to verify them—without adding extra model computation. By monitoring the draft‑token acceptance rate, SpecGuard detects when a target model shifts toward attacker‑controlled behavior while the draft model does not, signaling a backdoor trigger. The method reliably identifies a range of backdoor types, including stealthy attacks that bypass input‑level filters, across multiple model families.
By Rui Wen, Ahmed Salem, Andrew Paverd, Mark Russinovich, Zheng Li