arXiv Machine Learning

DiFA: Dual Evidence Fusion and Aggregation for Token-Level Text Anomaly Detection

arXiv AI
Aug 12

ProTAGAD: A Foundation Model for TAG Anomaly Detection with Decoupled Topological and Textual Prototypes

arXiv:2608. 10699v1 Announce Type: cross Abstract: Text-Attributed Graphs (TAGs), endowed with abundant textual content along with topological structures, have emerged as a versatile backbone for real-world anomaly detection spanning large language model security, social network moderation, and cyber threat identification.

By Ziyan Wang, Liwen Wu, Cheng Xie, Song Gao, Zhenli He, Xin Jin
arXiv Machine Learning
Sep 7

GLASS: Graph-Language Alignment with Spherical Scoring for Transferable Graph-Level Anomaly Detection

GLASS is a graph‑level anomaly detection framework that aligns graph and language representations on a unit hypersphere to achieve cross‑domain transferability. It constructs a Graph Descriptor Prompt to encode local, global, and semantic graph properties, and uses a multi‑slice soft cosine objective to unify graph and text embeddings. Anomaly scoring is performed via spherical density estimation with von Mises‑Fisher kernels, enabling zero‑shot detection and few‑shot adaptation across twelve benchmarks and three meta‑domains, outperforming recent GLAD baselines.

By Xudong Wang, Chris Ding, Tongxin Li, Jicong Fan
arXiv AI
Aug 28

Diff Mining: Logit Differences Reveal Finetuning Objectives

Diff Mining is a framework that identifies what a finetuned language model has learned by comparing its logits to those of its base model. It extracts per-context logit differences on a reference corpus and aggregates them into an interpretable token set using either a Top‑K frequency method or Non‑negative Matrix Factorization. The approach outperforms existing model‑diffing methods in domain detection and bias identification, and it requires only access to output logits, making it scalable to large models.

By Greg Kocher, Robert West, Cl\'ement Dumas, Julian Minder
arXiv Computation and Language
5d ago

Predicting Emerging Topics from Outliers: A Prospective Study of Weak Signals in Embedding Space

The study investigates whether documents initially classified as noise in embedding-based topic models can be identified as precursors to emerging topics. By labeling documents based on their future trajectories and measuring confidence across multiple embedding models, the authors find that anticipatory outliers are predictable at publication time, achieving an F1 score above 0.90 on high-consensus subsets and 0.76–0.80 in chronological evaluation. The predictive power largely stems from geometric features that capture each outlier’s position in embedding space.

By Evangelia Zve, Gauvain Bourgne, Jean-Gabriel Ganascia
arXiv AI
Sep 7

Adaptive Multi-Granularity Temporal Modeling for Weakly Supervised Video Anomaly Detection

The paper introduces an adaptive temporal modeling framework for weakly supervised video anomaly detection that addresses the limitations of rigid Multiple Instance Learning approaches. It presents a Temporal Refinement Module using dynamic positional encoding and a learnable class token to capture long‑range dependencies, and an Event Segmentation Module that identifies event boundaries via temporal discontinuity analysis to produce discriminative event‑level representations. An adaptive similarity‑based fusion strategy replaces fixed top‑k heuristics, dynamically integrating snippet‑level and event‑level anomaly scores into video‑level predictions, and the method outperforms state‑of‑the‑art baselines on two benchmarks.

By Changyi Li, Yu Xiao