arXiv AI By Tuan Nguyen, Sze Jue Yang, Khoa D. Doan, Chee Seng Chan, Kok-Seng Wong

FLAT: Revealing Hidden Latent-Conditioned Backdoor Failures in Federated Learning

Read the original on arXiv AI →

arXiv:2508. 04064v2 Announce Type: replace-cross Abstract: Horizontal federated learning (HFL) backdoor audits often summarize model behavior through clean accuracy (CA), mean attack success rate (ASR), or a single known-trigger test.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Jul 30

ToxScreen: Detecting Whether an LLM Has Been Poisoned

arXiv:2607. 26849v1 Announce Type: cross Abstract: As large language models (LLMs) are deployed in high-stakes domains, adversaries may poison training data to implant backdoors: hidden triggers that covertly manipulate model behavior at inference time.

By Anthony Hughes, Nicole Xing, Collin Francel, Andy Kim, Andrew Draganov
arXiv AI
Sep 24

Backdoors Leave Structural Traces: FedMAST for Backdoor Detection and Containment in Federated Learning

The paper introduces FedMAST, a Federated Multi‑Axis Structural Tracing defense designed to detect and contain backdoor attacks in federated learning. FedMAST evaluates client updates through complementary structural, spectral, and historical evidence, applying tiered filtering and round‑level containment. In experiments across six backdoor attacks, FedMAST consistently achieves lower attack success rates while preserving high main‑task accuracy.

By Srinivasan Subramanian, Kazi Aminul Islam, Md. Abdullah Al Hafiz Khan
Hugging Face Trending Papers
Jul 29

ToxScreen: Detecting Whether an LLM Has Been Poisoned

As large language models (LLMs) are deployed in high-stakes domains, adversaries may poison training data to implant backdoors: hidden triggers that covertly manipulate model behavior at inference time. We ask whether a defender can recover such a trigger under realistic affordances, namely white-box access to the weights and knowledge of the behavior of concern, but no training data, no trusted reference model, no knowledge of the trigger, and no certainty that the model is poisoned.

arXiv AI
Aug 28

LoRA as Oracle

The paper introduces a low‑rank auditing method called LoRA as Oracle, which fits a small adapter to a hypothesis and analyzes the geometry, energy, and alignment of the resulting update relative to frozen weights. This approach directly measures what a model has internalized, independent of its output behavior, enabling detection of backdoors that behavioral audits miss. By identifying and erasing malicious internalizations within the same low‑rank subspace, the method consistently detects target classes across multiple datasets and architectures while preserving clean accuracy and operating at far lower parameter and memory cost than full‑model baselines.

By Marco Arazzi, Antonino Nocera