arXiv Machine Learning

FAME: Failure-Aware Mixture-of-Experts for Message-Level Log Anomaly Detection

arXiv:2605. 22779v2 Announce Type: replace-cross Abstract: Production systems generate millions of log lines daily, yet most anomaly detectors operate at the session or window-level, flagging groups of lines rather than identifying the specific message responsible.

arXiv Machine Learning
5d ago

Towards Understanding LLM-Based Log Anomaly Detection: An Empirical Study of Performance, Efficiency, and Robustness

The paper investigates how different adaptation strategies, model architectures, parameter scales, and quantization settings influence the performance, efficiency, and robustness of large language models (LLMs) for log anomaly detection. Across three public log datasets, the study finds that adaptation strategies lead to significant performance variations, model scaling offers dataset‑dependent gains, and models with similar accuracy can differ markedly in computational cost. Low‑bit quantization largely preserves detection performance, and the authors also assess robustness to structural, semantic, and label noise at varying perturbation levels.

By Bin Li, Dongdong Wang, Siyang Lu
arXiv AI
Aug 19

Too Sure to Be Safe: Model Calibration for Reliable Log Anomaly Detection

The paper introduces LoRD, a lightweight post‑hoc calibration framework designed to improve confidence reliability in language‑model‑based log anomaly detectors. LoRD learns route‑specific reliability models from latent representations of correctly classified validation samples and uses reconstruction distances to estimate prediction reliability. By selectively recalibrating high‑risk predictions, LoRD reduces overconfident errors while maintaining strong anomaly detection performance across four large‑scale log benchmark datasets.

By Bin Li, Dongdong Wang, Siyang Lu
arXiv Machine Learning
Aug 21

From Noise to Signal: Improving Security Log Anomaly Detection Using LLMs with Endpoint-Specific Logs

arXiv:2608. 19938v1 Announce Type: cross Abstract: Existing approaches to anomalous behaviour log detection, such as Wazuh rely primarily on predefined detection rules, while statistical anomaly detection approaches such as OpenSearch identify deviations from previously observed behavioural patterns.

By Christopher Henshaw, Gour Karmakar
arXiv Machine Learning
Sep 3

TrajMind: Chaining Role-Specialized LoRAs for Fast-and-Slow Collective Trajectory Anomaly Diagnosis

TrajMind is a framework for diagnosing collective anomalies in urban trajectory data. It separates continuous screening from on-demand diagnosis, using a fast text-only path for alerts and a slow vision‑language path that chains role‑specialized LoRA adapters for detailed, evidence‑backed what‑who‑where‑when records. Experiments show the slow path outperforms baselines by over 15 percentage points in typing and 13 in localization, while the fast path cuts latency by 41% and retains high accuracy.

By Jiahao Wu, Zhenqun Yang, Chen Jason Zhang, Qing Li
arXiv AI
4d ago

MAADBench: The Refreshable Paradigm for Anomaly Detection in Multi-Agent Systems

MAADBench is a refreshable benchmark for anomaly detection in multi‑agent systems powered by large language models. It addresses the challenge of keeping benchmarks current by sampling and coupling generative tasks, generating trace data under configurable LLM backbones, and automatically providing deterministic step‑level labels. The authors evaluated 25 anomaly‑detection methods on 5,200 labeled traces, finding that existing approaches depend heavily on supervision, struggle with subtle MAS‑specific anomalies, and lack robustness across different LLM backbones.

By Lei Ma, Dennis Hofmann, Haowen Xu, Joshua DeOliveira, Peter VanNostrand, Lei Cao, Elke Rundensteiner
arXiv AI
Aug 28

Learning to Predict, Discover, and Reason in High-Dimensional Event Sequences

The paper proposes a new framework for automated fault diagnostics in modern vehicles by treating diagnostic trouble codes (DTCs) as a high‑dimensional language. It introduces Transformer‑based models for predictive maintenance, scalable causal discovery methods, and a multi‑agent system that automatically generates Boolean error‑pattern rules. The approach aims to replace costly manual grouping of DTCs with scalable, data‑driven techniques.

By Hugo Math