arXiv AI By William Smits

CRAFTIIF: Cross-Resolution Analytic Four-Type Interpretable Isolation Forest for Multivariate Time Series Anomaly Detection

Read the original on arXiv AI →

arXiv:2606. 13486v1 Announce Type: cross Abstract: Anomaly detection in multivariate time series is challenged by four structurally distinct anomaly types -- point (isolated spikes), distributional (level shifts), temporal (rhythm changes), and collective (inter-sensor correlation breakdowns) -- each requiring different feature representations.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Aug 26

A Hybrid Two-Stage Machine Learning Pipeline for Fault Detection and Classification in Power Transmission Systems

The paper introduces a hybrid two‑stage machine learning pipeline for fault detection and classification in high‑voltage transmission networks. Stage 1 uses an Isolation Forest anomaly detector combined with an optional supervised binary detector, while Stage 2 applies a Random Forest multiclass classifier only to samples flagged by Stage 1. Feature engineering maps six raw channels to eighteen features, including zero‑sequence symmetrical components, achieving end‑to‑end accuracies of 95.8 % on the TLFaultDataset and 97.25 % on an independent single‑point dataset, surpassing federated benchmarks without GPU or federated infrastructure.

By Sahil Manikshete, Atharva Gujarathi, Thanh Long Vu, Akhtar Hussain, Van-Hai Bui
arXiv AI
Sep 4

Differentiable Interval Bottlenecks for Interpretable Anomaly Detection in Numerical Data

DIFFINT is a reconstruction‑based anomaly detector that uses a differentiable autoencoder with a latent bottleneck composed of soft, axis‑aligned interval memberships. Each latent unit represents a human‑readable hyper‑rectangle in feature space, allowing the model to encode how strongly an instance falls inside each interval and to compute reconstruction error as the anomaly score. The method provides a certified lower bound on reconstruction error for points outside all active intervals, a suppression mechanism for sparse abnormalities, and a closed‑form, label‑free importance ranking for each (unit, feature) pair, achieving top performance on 48 ADBench benchmarks against 22 baselines.

By Lamine Diop, Marc Plantevit
arXiv AI
Sep 4

Witnesses Explain Anomalies

WAND is an unsupervised tabular anomaly detector that scores each point by how far its projection on unit‑sphere directions deviates from a sub‑Gaussian baseline. The directions that flag a point serve as its explanation, providing per‑feature attribution at no extra cost and recoverable via gradients. On 47 ADBench datasets, WAND matches or exceeds 16 baselines in ROC‑AUC while delivering more accurate, faithful explanations than post‑hoc SHAP, LIME, or ECOD, all with linear scoring time and a probe‑efficiency guarantee.

By Lamine Diop
Hugging Face Trending Papers
Sep 3

Differentiable Interval Bottlenecks for Interpretable Anomaly Detection in Numerical Data

DIFFINT is a reconstruction‑based anomaly detector that replaces the opaque latent bottleneck of a standard autoencoder with a set of soft, axis‑aligned interval memberships learned directly from raw numerical data. Each latent unit represents a human‑readable hyper‑rectangle, and an instance’s anomaly score is its reconstruction error weighted by how strongly it falls inside these intervals. The method provides a certified lower bound on reconstruction error for points outside all active intervals, a graded suppression mechanism for sparse anomalies, and a closed‑form, label‑free importance ranking for each (unit, feature) pair, achieving top performance on 48 ADBench benchmarks against 22 baselines. whyItMatters":"DIFFINT offers the first interpretable anomaly detector that maintains competitive performance while revealing which feature ranges drive each anomaly score, enabling practitioners to audit and understand model decisions without requiring anomaly labels."

arXiv Machine Learning
Sep 25

When Identical Rows Disagree: From Benchmark Identifiability to Replication-Robust Anomaly Detection

The paper investigates how repeated rows in released datasets—often treated as i.i.d. samples—introduce a hidden measurement layer that affects anomaly detection. It shows that identical rows can cap evaluation performance, make AUROC sensitive to replication, and bias detectors toward multiplicity size. The authors audit 690 OddBench datasets, find significant train-test overlap and label conflicts, and propose SCOUT, a support‑count orthogonalized detector that separates replication‑invariant evidence from exposure‑aware counts, achieving comparable or better AUROC while controlling false‑positive rates.

By Jie Deng