arXiv Machine Learning By Nils Loose, Jonas Sander, Felix M\"achtle, Thomas Eisenbarth

FloatDoor: Platform-Triggered Backdoors in LLMs

Read the original on arXiv Machine Learning →

arXiv:2606. 19535v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed in sensitive settings such as software engineering, where their outputs directly shape downstream artifacts.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
3d ago

Trusted Weights, Treacherous Optimizations? Optimization-Triggered Backdoor Attacks on LLMs

The paper investigates how inference optimization for large language models can introduce numerical inconsistencies that trigger hidden backdoors. It introduces two types of optimization‑triggered backdoors: the Input‑Specific Optimization Backdoor (ISOB) and the Universal Optimization Backdoor (UOB), the latter enabling a model to remain benign under normal execution but activate a backdoor when optimization is applied. Experiments on seven open‑source LLMs, across multiple tasks and optimization backends, show UOB can achieve up to 100% attack success while maintaining clean accuracy, and the authors propose three defenses that reduce the attack success rate to 0.02.

By Yifei Wang, Yida Yang, Tianlin Li, Xiaohan Zhang, Xiaoyu Zhang, Li Pan
arXiv Machine Learning
Sep 18

ALIBI: Adversarial Legitimacy Injection in Binary Input against LLM Malware Analyzers

The paper introduces ALIBI, a semantic cover story attack that injects a small, non-executed read‑only section into compiled binaries to mislead large language model (LLM) malware analyzers. By embedding a coherent but false security narrative, ALIBI can cause LLMs such as Gemini 2.5 Pro, GPT‑5.5 Pro, and Claude Opus 4.7 to downgrade or flip the verdicts of malicious samples. The attack also transfers to ELF binaries, and even a verification‑guided defense prompt only partially mitigates the effect, leaving a significant portion of malicious samples classified as benign.

By Hyeongjun Choi, Wonyoung Jung, Haehoon Seo, Sungyup Nam
arXiv Machine Learning
Aug 31

Not to Break, but to Attest: Adversarial Probes for Privacy-Preserving LLM Verification

The paper introduces a privacy‑preserving zk‑SNARK audit framework that uses adversarial‑style probes to detect logit drift between an approved large language model and a modified deployment. It offers three probe families—token‑based (black‑box), embedding‑based (gray‑box), and stress probes (partial white‑box)—allowing users to balance sensitivity, access, and cost. Experiments across LLM architectures and GPU platforms show token‑based probes achieve the highest mean sensitivity while remaining practical in a black‑box setting, with Groth16 proving times scaling modestly from 1.02 to 1.78 seconds and constant proof size.

By Cameron Wilding, Mina Shaker, Fatemeh Ganji