arXiv AI

Baszta: Data-Centric Fine-Tuning of a Polish Multi-Label Safety Classifier

The paper presents Baszta, a Polish multi‑label content‑safety classifier trained by fine‑tuning the 124M‑parameter allegro/herbert‑base‑cased model on five categories (hate, vulgarity, sexual content, crime, self‑harm) using a Focal + R‑Drop objective. In out‑of‑distribution evaluation on the Gadzi Język benchmark, Baszta achieves a small but statistically significant improvement in micro‑F1 over the Bielik Guard system, though the macro‑F1 advantage disappears when both models are properly tuned. The study also explores calibration techniques, showing that per‑category temperature scaling can recover performance lost by Platt scaling or isotonic regression, and discusses the trade‑offs between robust calibration and adversarial recall.

arXiv AI
Jul 28

Semalith v1.4: A Calibrated 184M Safety Classifier Achieving State-of-the-Art Prompt-Injection Detection at 44x Fewer Parameters than Llama-Guard-3-8B

arXiv:2607. 22545v1 Announce Type: cross Abstract: Deploying large language models in financial-services and agentic settings requires safety classifiers that simultaneously handle prompt injection, regulatory compliance, and general harm, a combination no existing open guardrail addresses in a single inference pass.

By Tejasvi C. Addagada
arXiv Machine Learning
Jun 25

How Reliable Is Your Jailbreak Judge? Calibration and Adversarial Robustness of Automated ASR Scoring

arXiv:2606. 25487v1 Announce Type: cross Abstract: Almost every paper on LLM jailbreaks and prompt injection reports an attack-success rate (ASR), and that number is assigned not by people but by an automated judge: either a safety classifier trained for the task, or a general chat model prompted to grade.

By Yang Gao (Veyon Solutions)
arXiv AI
Sep 1

Benchmark Contamination: A Taxonomy Organized by Defeated Mitigation

The paper introduces a new taxonomy for benchmark contamination that categorizes leakage by the mitigation it defeats—direct, derivative, temporal, distributional, and acquired—covering both training‑time and evaluation‑time scenarios. It proposes a four‑field disclosure protocol to record contamination status alongside benchmark scores, and provides a JSON schema, validator, and examples. An empirical study of 41 documents using a pre‑registered instrument shows limited reporting of contamination types and variable reliability, highlighting gaps in current disclosure practices.

By Johanna Angulo, V\'ictor Yeste, Hector Espinos-Morato
arXiv AI
Aug 25

A Reproducible, License-Aware Distillation Recipe for CPUDeployable Safety Classification

arXiv:2608.21570v1 Announce Type: new Abstract: Deploying a safety layer for large language models on commodity hardware is constrained by the guards available to do it: current open guard models hol...

By Edson Rodrigues da Cruz Filho, Paulo Ricardo Ferreira Neves, Paulo Henrique Eleuterio Falsetti, Jo\~ao Vitor Pavan, Ian Degaspari, Henrique Vieira Laturrague, Patrick Vieira Laturrague, Guilherme Nielsen Dias, Marccello Wilson Perez Berto, Gustavo Voltani Von Atzingen