The paper identifies a problem in multi‑view anomaly detection called cross‑view information leakage, where fusing multiple inspection views can cause normal features to mask anomalies during reconstruction. To address this, the authors propose GLAD, a framework that uses a Global‑Local Attention Driven approach, combining vision foundation model features with two fusion modules: Multi‑view Merging Attention for local, weighted fusion and Object‑Guided Attention for global context aggregation. Experiments on Real‑IAD and MANTA‑Tiny demonstrate that GLAD outperforms existing methods across various metrics, underscoring the importance of restricting information flow to preserve the reconstruction gap.
By Shang-Fu Chen, Kuan-Chuan Peng, Jhih-Ciang Wu, Wen-Huang Cheng, Kai-Lung Hua
arXiv:2603. 26842v3 Announce Type: replace-cross Abstract: Time series anomaly detection (TSAD) is essential for maintaining the reliability and security of IoT-enabled service systems.
By PengYu Chen, Shang Wan, Xiaohou Shi, Yuan Chang, Yan Sun, Sajal K. Das
arXiv:2512. 22179v3 Announce Type: replace Abstract: Detecting previously unseen attacks remains a major challenge for machine learning-based intrusion detection systems.
By Rajeeb Thapa Chhetri, Saurab Thapa, Avinash Kumar, Zhixiong Chen
arXiv:2608. 06876v1 Announce Type: cross Abstract: In the era of Industrial Internet of Things (IIoT) and Cyber-Physical Systems (CPS), Federated Learning (FL) offers a promising decentralized intelligence paradigm for Video Anomaly Recognition (VAR).
By Ghani Haider, Majid Kundroo, Boyun Eom, Dong Hwan Park, Chen Chen, Taehong Kim
arXiv:2602. 08638v2 Announce Type: replace-cross Abstract: As a fundamental data mining task, unsupervised time series anomaly detection (TSAD) aims to build a model for identifying abnormal timestamps without assuming the availability of annotations.
By Dezheng Wang, Tong Chen, Guansong Pang, Congyan Chen, Shihua Li, Hongzhi Yin
The paper introduces NC‑TFAD, a task‑free continual anomaly detection framework that leverages neural‑collapse geometry to learn from non‑stationary data streams without task boundaries. It freezes a pretrained backbone, aligns streaming features to a simplex Equiangular Tight Frame prototype space, and uses synthetic anomaly anchors, inter‑ and intra‑class regularization, and a Focal Neural Collapse Contrastive loss to stabilize representations and enhance normal‑anomaly separability. A normal‑patch‑prototype‑guided localization branch generates calibrated anomaly heatmaps, and extensive experiments on MVTec AD and VisA demonstrate that NC‑TFAD outperforms existing task‑free continual learning and unified anomaly detection baselines in both image‑level detection and pixel‑level localization.
By Xiaotong Kong, Chaoyang Song, Ziai Zhou, Jinxia Zhang, Kanjian Zhang, Haikun Wei
TrajMind is a framework for diagnosing collective anomalies in urban trajectory data. It separates continuous screening from on-demand diagnosis, using a fast text-only path for alerts and a slow vision‑language path that chains role‑specialized LoRA adapters for detailed, evidence‑backed what‑who‑where‑when records. Experiments show the slow path outperforms baselines by over 15 percentage points in typing and 13 in localization, while the fast path cuts latency by 41% and retains high accuracy.
By Jiahao Wu, Zhenqun Yang, Chen Jason Zhang, Qing Li
GeoMAD is a multi‑view anomaly detection framework that fuses multiple camera viewpoints while maintaining geometric awareness and scalability to multi‑class industrial settings. It introduces a Cross‑view Deformable Fusion Module (CDFM) that learns view‑pair‑specific sampling offsets on 2D feature maps, enabling hierarchical cross‑view correspondence without camera calibration or voxel construction. Additionally, Distributional View Alignment (DVA) provides a self‑supervised loss that aligns bottleneck distributions across views, ensuring global consistency without pixel‑level correspondence. Together, CDFM and DVA achieve geometry‑aware, distribution‑consistent fusion and demonstrate strong detection and localization performance on Real‑IAD and MANTA‑Tiny datasets.
By Shang-Fu Chen, Jhih-Ciang Wu, Kuan-Chuan Peng, Wen-Huang Cheng, Kai-Lung Hua
The paper introduces NC‑TFAD, a task‑free continual anomaly detection framework that leverages neural‑collapse geometry to handle non‑stationary data streams in industrial visual inspection. It freezes a pretrained backbone, aligns streaming features to a simplex Equiangular Tight Frame prototype space, and uses synthetic anomaly anchors, inter‑ and intra‑class regularization, and a Focal Neural Collapse Contrastive loss to stabilize representations and enhance normal‑anomaly separability. A normal‑patch‑prototype‑guided localization branch generates calibrated anomaly heatmaps without pixel‑level annotations, and experiments on MVTec AD and VisA demonstrate that NC‑TFAD outperforms existing task‑free continual learning and unified anomaly detection baselines in both image‑level detection and pixel‑level localization.
While recent advancements in anomaly detection have demonstrated the efficacy of CNN- and Transformer-based approaches, these architectures face inherent limitations: CNNs struggle to capture long-range dependencies, whereas Transformers suffer from quadratic computational complexity. Consequently, Mamba-based architectures have attracted considerable attention, as they successfully combine superior long-range dependency modeling with linear computational complexity.
The paper introduces the concept of protocol divergence, showing that identical nominal missing rates can lead to vastly different learning regimes in incomplete multi‑view clustering. It critiques existing evaluation practices that ignore observation structure and proposes CRAFT, a train‑once framework that fuses observed views with mask‑aware attention, enabling efficient deployment across multiple missing‑view protocols. Experiments on CUB, MultiFashion, and other benchmarks demonstrate CRAFT’s superior performance and significant computational savings through checkpoint reuse.
By Haolu Liu, Xiyue Wang, Xuanting Xie, Liangjian Wen, Zhao Kang
The paper introduces TITAnD, a Trajectory Image Transformer that converts dense and sparse GPS trajectories into a Hyperspectral Trajectory Image (HTI) and applies vision-based classification and segmentation for anomaly detection. It employs a Cyclic Factorized Transformer (CFT) that splits attention along within-day and across-day axes, drastically reducing computational cost and enabling multi-month analysis. Empirical results show TITAnD outperforms existing sparse and dense benchmarks, achieving higher AUC-PR and faster inference than comparable Transformers.
By Md Awsafur Rahman, Chandrakanth Gudavalli, Hardik Prajapati, B. S. Manjunath