arXiv Machine Learning

NLLog: Lightweight, Explainable SOC Anomaly Detection via Log-to-Language Rewriting

arXiv:2606. 04957v1 Announce Type: cross Abstract: System-generated logs underpin security monitoring, yet their rigid template-based format hinders both automated analysis and human comprehension.

arXiv Machine Learning
Jul 21

FAME: Failure-Aware Mixture-of-Experts for Message-Level Log Anomaly Detection

arXiv:2605. 22779v2 Announce Type: replace-cross Abstract: Production systems generate millions of log lines daily, yet most anomaly detectors operate at the session or window-level, flagging groups of lines rather than identifying the specific message responsible.

By Huanchi Wang, Zihang Huang, Yifang Tian, Kristina Dzeparoska, Hans-Arno Jacobsen, Alberto Leon-Garcia
arXiv AI
Aug 19

Too Sure to Be Safe: Model Calibration for Reliable Log Anomaly Detection

The paper introduces LoRD, a lightweight post‑hoc calibration framework designed to improve confidence reliability in language‑model‑based log anomaly detectors. LoRD learns route‑specific reliability models from latent representations of correctly classified validation samples and uses reconstruction distances to estimate prediction reliability. By selectively recalibrating high‑risk predictions, LoRD reduces overconfident errors while maintaining strong anomaly detection performance across four large‑scale log benchmark datasets.

By Bin Li, Dongdong Wang, Siyang Lu
arXiv Machine Learning
5d ago

Towards Understanding LLM-Based Log Anomaly Detection: An Empirical Study of Performance, Efficiency, and Robustness

The paper investigates how different adaptation strategies, model architectures, parameter scales, and quantization settings influence the performance, efficiency, and robustness of large language models (LLMs) for log anomaly detection. Across three public log datasets, the study finds that adaptation strategies lead to significant performance variations, model scaling offers dataset‑dependent gains, and models with similar accuracy can differ markedly in computational cost. Low‑bit quantization largely preserves detection performance, and the authors also assess robustness to structural, semantic, and label noise at varying perturbation levels.

By Bin Li, Dongdong Wang, Siyang Lu
arXiv Machine Learning
Aug 28

Stack Trace-Based Crash Deduplication with Transformer Adaptation

Stack Trace-Based Crash Deduplication with Transformer Adaptation introduces dedupT, a transformer‑based method that models entire stack traces instead of isolated frames. The approach first fine‑tunes a pretrained language model on stack traces and then trains a fully‑connected network to rank duplicate crashes. Experiments on four public datasets show dedupT improves Mean Reciprocal Rank by over 15% versus the best deep‑learning baseline and up to 10% over traditional methods, while also achieving higher ROC‑AUC for unique crash detection.

By Md Afif Al Mamun, Gias Uddin, Lan Xia, Longyu Zhang