arXiv AI

Product-Aware Deep Autoencoders for Robust Process Monitoring in Multi-Product Cyber-Physical Systems

arXiv:2606. 00052v1 Announce Type: new Abstract: As Industry 4.

arXiv Computer Vision
Sep 4

ISP-AD: A Large-Scale Real-World Dataset for Advancing Industrial Anomaly Detection with Synthetic and Real Defects

The paper introduces ISP-AD, the largest publicly available industrial anomaly detection dataset, featuring both synthetic and real defects from a factory floor. It focuses on challenging, small, weakly contrasted surface defects within highly variable structured patterns, addressing the bias of existing datasets toward ideal imaging conditions. Experiments demonstrate that even a small amount of weakly labeled real defects improves model generalization and that synthetic defects can serve as a useful cold‑start baseline for scalable training.

By Paul J. Krassnig, Dieter P. Gruber
arXiv Machine Learning
Aug 11

From Benchmark Performance to Tool Deployment: Human-in-the-Loop Anomaly Detection

arXiv:2608. 07770v1 Announce Type: new Abstract: Automated anomaly detection methods often report strong performance on curated academic benchmarks, but their behavior under real-world industrial conditions is less clear.

By Mike Szklarzewski, CJ George, Gavin Smithson, Christopher Stokes, Dakota Fulp, William M. Jones, Benjamin Wynn, Alexander Ur, Agit Yesiloz, Clint Kallenbach, Mark Swartz, Nathan DeBardeleben, Sharmistha Chakrabarti
arXiv Machine Learning
Aug 20

Online Conformal Anomaly Detection with Prediction-Powered Data Acquisition

Online Conformal Anomaly Detection with Prediction-Powered Data Acquisition introduces C-PP-COAD, a framework that uses synthetic calibration data to reduce reliance on real-world calibration while maintaining assumption-free false discovery rate control. The method wraps any anomaly detection algorithm, converting its scores into conformal p-values for online testing. Experiments on synthetic and real datasets—including thyroid dysfunction, O‑RAN conflict, 5G intrusion, and UE throughput degradation—show that C-PP-COAD preserves FDR guarantees while significantly cutting the need for real calibration data.

By Amirmohammad Farzaneh, Osvaldo Simeone
arXiv Machine Learning
Sep 21

Identifying Security Platform Product Abuse with Machine Learning

arXiv:2609.21303v1 Announce Type: cross Abstract: Product abuse is an individually rare, but growing, problem across the SaaS industry. Highly sophisticated threat actors can misuse security platform...

By Shaefer Drew, Michael Brautbar, Paul Knight, Edward Raff, Lana Peric-McDermott, Simran Sarin, Nickolas Machado, Hanna Albright, Vitaly Zaytsev
arXiv AI
Sep 10

Towards a Resilience-Theoretic Foundation for Adversarial Robustness in Industrial Control System Anomaly Detection

The paper argues that adversarial robustness of anomaly detectors in industrial control systems (ICS) is a specific form of system resilience. It maps resilience concepts—disturbance class, absorption capacity, recovery trajectory, and degradation function—to adversarial machine learning, deriving a compositional resilience bound that identifies the coupling‑adjusted absorption capacity of nodes along an attack path as the key constraint. Empirical tests on the BATADAL water distribution benchmark reveal operationally significant effects, such as absorption‑degradation divergence under adversarial training and a paradox where hardening the most vulnerable node alone can reduce overall resilience.

By Branka Stojanovi\'c, Andreas Flatscher, Michael Somma
arXiv AI
Sep 18

Local Sparsity Enables Unsupervised LLM Safety Detection

The paper proposes a new unsupervised safety detection method for large language models that relies on anomaly detection rather than supervised training on unsafe data. By leveraging local sparsity in a linear representation space obtained via a sparse autoencoder, the authors develop a framework for locally masked SAE-based anomaly detection, providing theoretical support and empirical validation across multiple architectures and datasets. When calibrated with only 1% out-of-distribution data, the method achieves near‑optimal performance while using just 1–2% of SAE neurons for computation.

By Xin Chen, Gil Kur, Alexander Shevchenko, Andreas Krause