arXiv:2608. 14089v1 Announce Type: new Abstract: Safety classifiers deployed with large language models often fail for two reasons: their decisions reflect the policy learned during training rather than the deployer's desired policy, and their performance degrades as deployment traffic evolves.
By Thiago Sandoval, Ufuk Topcu
arXiv:2607. 26574v1 Announce Type: cross Abstract: Safety classifiers ("guards") are the dominant black-box defense for vision-language models, yet they judge an input's surface form, not its meaning: a harmful request re-encoded as set theory, formal logic, a rare language, code, or an image of text slips past a guard that would block it in plain language -- the decode gap.
By Haoyu Zhang, Zhuoxi Wang, Shibo Zheng, Zijian Xiao, Xiangchen Guan, Mohammad Zandsalimy, Shanu Sushmita
arXiv:2607. 01679v1 Announce Type: cross Abstract: Adversarial attacks on cybersecurity classifiers pose a dual threat: degrading predictions and destabilising the SHAP-based explanations that security analysts rely on to understand and triage alerts.
By Mona Rajhans, Vishal Khawarey
arXiv:2512. 19011v3 Announce Type: replace-cross Abstract: Safety classifiers that screen LLM inputs for jailbreak attempts have become standard deployment components, yet almost all production systems rely on GPU-based models: fine-tuned transformers and LLM-as-a-judge pipelines.
By Vasudev Majhi, Dhruv Gupta, Advait Singh, Matthew Barker, Dhruv Kumar
arXiv:2609.13534v1 Announce Type: new
Abstract: We identify \textbf{Harmfulness Propagation Dynamics (HPD)}: for harmful prompts, the projection of the last-token hidden state onto a learned harm dir...
By Noor Islam S. Mohammad, Ulu\u{g} Bayaz{\i}t
arXiv:2605. 24696v2 Announce Type: replace-cross Abstract: Streaming intrusion detection systems must process flows continuously under bounded memory, yet most leave alerting-threshold selection as a post-hoc tuning problem incompatible with production, where operators commit in advance to alert budgets, misclassification costs, and Service Level Objectives.
By Michel A. Youssef
arXiv:2607. 12792v1 Announce Type: cross Abstract: Jailbreak-robustness research typically evaluates safety through generated responses using an LLM-as-judge approach.
By Roman Prosvirnin, Victor Minchenkov, Alexey Soldatov, Vladimir Bashun
arXiv:2602. 14161v2 Announce Type: replace Abstract: Detecting prompt injection, jailbreak attacks, and harmful requests is critical for deploying LLM-based agents safely, yet current evaluation practices in this literature overestimate generalization.
By Max Fomin
arXiv:2505. 04608v5 Announce Type: replace-cross Abstract: Responsibly deploying artificial intelligence (AI) / machine learning (ML) systems in high-stakes settings arguably requires not only proof of system reliability, but also continual, post-deployment monitoring to quickly detect and address any unsafe behavior.
By Drew Prinster, Xing Han, Anqi Liu, Suchi Saria
The paper introduces FedMAST, a Federated Multi‑Axis Structural Tracing defense designed to detect and contain backdoor attacks in federated learning. FedMAST evaluates client updates through complementary structural, spectral, and historical evidence, applying tiered filtering and round‑level containment. In experiments across six backdoor attacks, FedMAST consistently achieves lower attack success rates while preserving high main‑task accuracy.
By Srinivasan Subramanian, Kazi Aminul Islam, Md. Abdullah Al Hafiz Khan
arXiv:2505. 14300v2 Announce Type: replace Abstract: White-box monitoring is increasingly adopted as an auditing tool as Large Language Models (LLMs) are deployed in daily operations to ensure safe model behavior.
By Maheep Chaudhary, Fazl Barez
CS-WCP introduces confidence‑set weighted conformal prediction to provide robust prediction sets for large‑language‑model judges when deployment traffic shifts the prevalence of task or policy groups. By constructing simultaneous exact intervals for source and target group masses and taking the union over all compatible ratio vectors, CS‑WCP achieves high coverage (mean 0.973) with few failures across 336 constructed traffic shifts, outperforming standard source conformal prediction. The method offers an auditable coverage safeguard under uncertain mixture weights, focusing on conservative tail protection rather than tighter set sizes.
By Ibne Farabi Shihab, Fariya Afrin