arXiv:2606. 20502v1 Announce Type: cross Abstract: Whether LLMs scoring well on vulnerability benchmarks genuinely reason about security or merely pattern-match on contaminated data remains unresolved.
By Arastoo Zibaeirad, Marco Vieira
The paper introduces TRIM, a black‑box defense for backdoor attacks in computer vision models. TRIM identifies and removes malicious trigger regions at inference time using region‑based segmentation, adaptive trigger discovery via inpainting and diffusion, and selective purification, without needing model internals, training data, or clean samples. Experiments on various datasets and trigger types show TRIM reduces attack success rates to as low as 1.16% while maintaining high clean accuracy.
By Ahmed Abdelnaby, Mohamed Elmahallawy
The paper investigates how inference optimization for large language models can introduce numerical inconsistencies that trigger hidden backdoors. It introduces two types of optimization‑triggered backdoors: the Input‑Specific Optimization Backdoor (ISOB) and the Universal Optimization Backdoor (UOB), the latter enabling a model to remain benign under normal execution but activate a backdoor when optimization is applied. Experiments on seven open‑source LLMs, across multiple tasks and optimization backends, show UOB can achieve up to 100% attack success while maintaining clean accuracy, and the authors propose three defenses that reduce the attack success rate to 0.02.
By Yifei Wang, Yida Yang, Tianlin Li, Xiaohan Zhang, Xiaoyu Zhang, Li Pan
arXiv:2606. 01695v1 Announce Type: new Abstract: Adversaries can implant latent harmful behavior by poisoning as few as 1% of fine-tuning examples.
By Swapnil Parekh
arXiv:2602. 14161v2 Announce Type: replace Abstract: Detecting prompt injection, jailbreak attacks, and harmful requests is critical for deploying LLM-based agents safely, yet current evaluation practices in this literature overestimate generalization.
By Max Fomin
SpecGuard is an inference‑time backdoor detector that leverages speculative decoding—using a small draft model to propose tokens and a target model to verify them—without adding extra model computation. By monitoring the draft‑token acceptance rate, SpecGuard detects when a target model shifts toward attacker‑controlled behavior while the draft model does not, signaling a backdoor trigger. The method reliably identifies a range of backdoor types, including stealthy attacks that bypass input‑level filters, across multiple model families.
By Rui Wen, Ahmed Salem, Andrew Paverd, Mark Russinovich, Zheng Li
arXiv:2608.24354v1 Announce Type: cross
Abstract: MLLMs are increasingly deployed in user-facing applications, yet they inherit backdoor risks from the pipelines used to construct them: triggers may...
By Jiali Wei, Ming Fan, Mingkun Zhang, Haoyu Wang, Jun Sun, Guoheng Sun, Xiaoning Ren, Haijun Wang, Ting Liu
The paper introduces a low‑rank auditing method called LoRA as Oracle, which fits a small adapter to a hypothesis and analyzes the geometry, energy, and alignment of the resulting update relative to frozen weights. This approach directly measures what a model has internalized, independent of its output behavior, enabling detection of backdoors that behavioral audits miss. By identifying and erasing malicious internalizations within the same low‑rank subspace, the method consistently detects target classes across multiple datasets and architectures while preserving clean accuracy and operating at far lower parameter and memory cost than full‑model baselines.
By Marco Arazzi, Antonino Nocera
The paper introduces Dynamic DAE Guardrails (DSG), a method that uses Dynamic Sparse Autoencoders to perform precision unlearning in large language models. DSG leverages principled feature selection and a dynamic classifier to target activation-based unlearning, outperforming existing gradient‑based methods in terms of computational efficiency, stability, sequential unlearning, resistance to relearning attacks, data efficiency, and interpretability.
By Aashiq Muhamed, Jacopo Bonato, Mona Diab, Virginia Smith
arXiv:2609. 12591v1 Announce Type: new Abstract: Foundation models are increasingly adapted through fine-tuning, model editing, and alignment procedures while retaining previously acquired capabilities.
By Hendrik Droste, Christian Medeiros Adriano, Kathrin Korte, Holger Giese
arXiv:2608. 02271v1 Announce Type: new Abstract: Parameter-Efficient Fine-tuned (PEFT) models are frequently downloaded from open repositories by practitioners.
By Nicola Pitzalis, Donald Shenaj, Giacomo Cignoni, Andrea Cossu, Davide Bacciu, Antonio Carta
arXiv:2508. 04064v2 Announce Type: replace-cross Abstract: Horizontal federated learning (HFL) backdoor audits often summarize model behavior through clean accuracy (CA), mean attack success rate (ASR), or a single known-trigger test.
By Tuan Nguyen, Sze Jue Yang, Khoa D. Doan, Chee Seng Chan, Kok-Seng Wong