arXiv Machine Learning

Online Shift Detection and Conformal Adaptation for Deployed Safety Classifiers

arXiv:2606. 11949v1 Announce Type: new Abstract: We present an online monitoring system for distributional shift in deployed safety classifiers, using calibrated sequential statistics to detect when a classifier has moved out of distribution.

arXiv Machine Learning
Jul 30

Recover, Decode, Reguard: Guard-Agnostic Defense Amplification againstEncoded VLM Jailbreaks

arXiv:2607. 26574v1 Announce Type: cross Abstract: Safety classifiers ("guards") are the dominant black-box defense for vision-language models, yet they judge an input's surface form, not its meaning: a harmful request re-encoded as set theory, formal logic, a rare language, code, or an image of text slips past a guard that would block it in plain language -- the decode gap.

By Haoyu Zhang, Zhuoxi Wang, Shibo Zheng, Zijian Xiao, Xiangchen Guan, Mohammad Zandsalimy, Shanu Sushmita
arXiv AI
Jun 16

Do You Really Need a GPU to Guard Your LLM? CPU-Class Classifiers and Multi-Stage Pipelines for Safety Enforcement at Scale

arXiv:2512. 19011v3 Announce Type: replace-cross Abstract: Safety classifiers that screen LLM inputs for jailbreak attempts have become standard deployment components, yet almost all production systems rely on GPU-based models: fine-tuned transformers and LLM-as-a-judge pipelines.

By Vasudev Majhi, Dhruv Gupta, Advait Singh, Matthew Barker, Dhruv Kumar
arXiv Machine Learning
Jun 26

CALIBURN: Operationally Calibrated Streaming Intrusion Detection with Regime-Dependent Conformal Risk Control

arXiv:2605. 24696v2 Announce Type: replace-cross Abstract: Streaming intrusion detection systems must process flows continuously under bounded memory, yet most leave alerting-threshold selection as a post-hoc tuning problem incompatible with production, where operators commit in advance to alert budgets, misclassification costs, and Service Level Objectives.

By Michel A. Youssef
arXiv AI
Sep 24

Backdoors Leave Structural Traces: FedMAST for Backdoor Detection and Containment in Federated Learning

The paper introduces FedMAST, a Federated Multi‑Axis Structural Tracing defense designed to detect and contain backdoor attacks in federated learning. FedMAST evaluates client updates through complementary structural, spectral, and historical evidence, applying tiered filtering and round‑level containment. In experiments across six backdoor attacks, FedMAST consistently achieves lower attack success rates while preserving high main‑task accuracy.

By Srinivasan Subramanian, Kazi Aminul Islam, Md. Abdullah Al Hafiz Khan
arXiv Machine Learning
Sep 24

CS-WCP: Robust Conformal Sets for LLM-Judge Traffic Shifts with Uncertain Group Proportions

CS-WCP introduces confidence‑set weighted conformal prediction to provide robust prediction sets for large‑language‑model judges when deployment traffic shifts the prevalence of task or policy groups. By constructing simultaneous exact intervals for source and target group masses and taking the union over all compatible ratio vectors, CS‑WCP achieves high coverage (mean 0.973) with few failures across 336 constructed traffic shifts, outperforming standard source conformal prediction. The method offers an auditable coverage safeguard under uncertain mixture weights, focusing on conservative tail protection rather than tighter set sizes.

By Ibne Farabi Shihab, Fariya Afrin