arXiv Machine Learning By Fatih Deniz, Yazan Boshmaf, Issa Khalil

SSP-Bench: A Hybrid Data Generation Framework for Safety, Security, and Privacy Evaluation

Read the original on arXiv Machine Learning →

SSP-Bench is a dynamic benchmarking framework designed to evaluate large language models on safety, security, and privacy (SSP) by generating evaluation instances on demand while maintaining domain consistency. It ensures label validity through external sources, enforces scope with service-specific validation, and calibrates difficulty using a multi-model steering panel, framing benchmark construction as a multi-objective optimization over difficulty, separability, novelty, and diversity. Across 24 models and four SSP services, SSP-Bench exposes systematic failures of static benchmarks, such as near-zero correlation in safety rankings, strong safety–over-refusal coupling, and hidden within-family regressions.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Aug 19

Benchmarking the Benchmarks: Evaluating Automated Safety Benchmarks for Small Language Models

The paper investigates whether existing AI safety benchmarks, designed for large language models, are suitable for evaluating small language models (SLMs). By testing five benchmark suites on 26 open‑source SLMs with a unified scoring rubric, the authors find that ambiguous judgments dominate, especially for complex prompts and certain architectures. This ambiguity, linked to factors like lexical density and output perplexity, undermines the reliability of aggregate leaderboards and reveals a confound between model capability and perceived safety.

By Nyamtulla Shaik, Fengjun Li, Bo Luo
arXiv AI
Sep 10

Benchmark Scores Are Pipeline-Dependent: A Reliability Audit of Cybersecurity LLM Benchmarks

The paper examines how the scores of cybersecurity large language model (LLM) benchmarks vary depending on the evaluation pipeline used. By auditing eight benchmarks across ten different LLMs, the authors uncover 15 systematic failure modes and demonstrate that a single pipeline choice can shift a model’s score by over 80 percentage points, significantly altering rankings. They also show that even semantically similar tasks can produce different model rankings due to incompatible evaluation conventions, and that standardizing pipelines can move most models by at least three ranks on at least one benchmark.

By Aymene Berriche, Cathrine Shalby, Mohannad Alhanahnah, Yazan Boshmaf