arXiv Machine Learning

MINES: Explainable Anomaly Detection through Web API Invariant Inference

arXiv:2512. 06906v2 Announce Type: replace-cross Abstract: Detecting the anomalies of web applications, important infrastructures for running modern companies and governments, is crucial for providing reliable web services.

arXiv Machine Learning
Jul 21

FAME: Failure-Aware Mixture-of-Experts for Message-Level Log Anomaly Detection

arXiv:2605. 22779v2 Announce Type: replace-cross Abstract: Production systems generate millions of log lines daily, yet most anomaly detectors operate at the session or window-level, flagging groups of lines rather than identifying the specific message responsible.

By Huanchi Wang, Zihang Huang, Yifang Tian, Kristina Dzeparoska, Hans-Arno Jacobsen, Alberto Leon-Garcia
arXiv AI
4d ago

MAADBench: The Refreshable Paradigm for Anomaly Detection in Multi-Agent Systems

MAADBench is a refreshable benchmark for anomaly detection in multi‑agent systems powered by large language models. It addresses the challenge of keeping benchmarks current by sampling and coupling generative tasks, generating trace data under configurable LLM backbones, and automatically providing deterministic step‑level labels. The authors evaluated 25 anomaly‑detection methods on 5,200 labeled traces, finding that existing approaches depend heavily on supervision, struggle with subtle MAS‑specific anomalies, and lack robustness across different LLM backbones.

By Lei Ma, Dennis Hofmann, Haowen Xu, Joshua DeOliveira, Peter VanNostrand, Lei Cao, Elke Rundensteiner
arXiv Machine Learning
Aug 21

From Noise to Signal: Improving Security Log Anomaly Detection Using LLMs with Endpoint-Specific Logs

arXiv:2608. 19938v1 Announce Type: cross Abstract: Existing approaches to anomalous behaviour log detection, such as Wazuh rely primarily on predefined detection rules, while statistical anomaly detection approaches such as OpenSearch identify deviations from previously observed behavioural patterns.

By Christopher Henshaw, Gour Karmakar
arXiv Computation and Language
Aug 28

An Embarrassingly Simple Detector for Model Extraction Attacks in Large Language Model API Traffic

The paper presents a simple detector for model extraction attacks on large language model APIs. It frames detection as a benign‑calibrated traffic‑window distribution test, embedding queries into a semantic space and using maximum mean discrepancy (MMD) to compare against historical benign traffic. Evaluated on fourteen attacker‑normal query pairs across four extraction scenarios, MMD achieves near‑perfect true‑positive rates while maintaining a 0.3% false‑positive rate, outperforming several existing baselines.

By Shuze Liu, Qianwen Guo, Yushun Dong
arXiv Machine Learning
5d ago

Towards Understanding LLM-Based Log Anomaly Detection: An Empirical Study of Performance, Efficiency, and Robustness

The paper investigates how different adaptation strategies, model architectures, parameter scales, and quantization settings influence the performance, efficiency, and robustness of large language models (LLMs) for log anomaly detection. Across three public log datasets, the study finds that adaptation strategies lead to significant performance variations, model scaling offers dataset‑dependent gains, and models with similar accuracy can differ markedly in computational cost. Low‑bit quantization largely preserves detection performance, and the authors also assess robustness to structural, semantic, and label noise at varying perturbation levels.

By Bin Li, Dongdong Wang, Siyang Lu
arXiv Machine Learning
Aug 4

How Benchmarks and Evaluation Protocols Shape Conclusions in Provenance-Based Intrusion Detection

arXiv:2608. 01454v1 Announce Type: cross Abstract: Provenance-based intrusion detection systems (PIDS) frequently report strong performance, but the conclusions drawn from these results can be highly sensitive to benchmarking choices and evaluation protocols.

By Lorenzo Guerra, Thomas Chapuis, Guillaume Duc, Pavlo Mozharovskyi, Van-Tam Nguyen