arXiv AI

EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures

arXiv:2606. 30219v1 Announce Type: new Abstract: LLM evaluation and AI safety face a shared measurement problem: benchmark scores, reward-model signals, and reported safety metrics can improve while the latent properties they are meant to represent remain difficult to verify.

arXiv AI
Jul 31

Adversarial Pragmatics for AI Safety Evaluation: A Diagnostic Framework and Seed Benchmark for Language-Mediated Control

arXiv:2607. 01153v3 Announce Type: replace-cross Abstract: Safety evaluations for language models increasingly depend on judgments about ambiguous natural-language behaviour: whether a model followed an instruction, refused appropriately, complied with a policy, or misreported progress in an agentic task.

By Brett Reynolds
arXiv AI
2d ago

A Deterministic and Auditable AI Security Risk Assessment Framework with ATLAS Aligned Executable Rules and Formal Verification

The paper introduces a deterministic AI security risk assessment framework that transforms diverse engineering artefacts into a standardized Control ID taxonomy scored on a four‑level ordinal scale. It compiles technique‑level predicates from a fixed MITRE ATLAS snapshot, linking each control to mitigation and producing traceable feasibility and impact outputs. The framework is formally verified for boundedness, totality, consistency, and monotonicity, and is evaluated on five open‑source AI projects, showing that strengthened controls lower feasibility scores while residual risks persist when core controls are missing.

By Yixuan Huang (University of Southampton, Southampton, UK), Basel Halak (University of Southampton, Southampton, UK), Boojoong Kang (University of Southampton, Southampton, UK)
arXiv AI
Jul 2

Adversarial Pragmatics for AI Safety Evaluation: A Benchmark for Instruction Conflict, Embedded Commands, and Policy Ambiguity

arXiv:2607. 01153v1 Announce Type: cross Abstract: Safety evaluations for language models increasingly depend on judgments about ambiguous natural-language behaviour: whether a model has followed an instruction, refused appropriately, complied with a policy, resisted an embedded command, or misreported progress in an agentic task.

By Brett Reynolds
arXiv AI
3d ago

Trust Is Not a Score: Runtime Assurance Contracts for High-Risk AI Agents

The paper introduces Runtime Assurance Contracts (RAC) as a formal policy framework for high‑risk AI agents, addressing the "assurance‑transition gap" by binding autonomy boundaries, component eligibility, evidence state, transition policy, human‑review capacity, and non‑compensatory gates. RAC allows soft metrics to influence routing while mandating retries, switches, escalations, deferrals, or stops when mandatory gates fail or are unknown, ensuring aggregate performance cannot alone authorize action. The authors define the contract, evidence record, permission rule, and five invariants, and evaluate RAC through deterministic failure‑injection studies, hand‑authored traces, and a prospective synthetic holdout, comparing it to score‑only and restricted protocol baselines.

By Serhii Zabolotnii
arXiv AI
Sep 2

Validity-Aware Jailbreak Evaluation for Large Language Models

The paper introduces SEAV, a verification‑centric framework for evaluating jailbreak attempts against large language models. SEAV decomposes responses into ordered steps and checks both validity and correctness using LLM‑as‑a‑judge and retrieval‑grounded verification. The method reduces false positives by 14.9 percentage points on a strategic‑dishonesty diagnostic and reclassifies 22.1–51.0% of previously successful jailbreaks as invalid across multiple benchmarks.

By Qilong Wu, Sahil Wadhwa, Pranab Mohanty, Giri Iyengar, Varun Chandrasekaran
arXiv Machine Learning
1d ago

The Alignment Flywheel: A Governance-Centric Hybrid MAS for Architecture-Agnostic Safety

The paper introduces the Alignment Flywheel, a governance‑centric hybrid multi‑agent system (MAS) that separates decision generation from safety governance. It defines a Proposer that generates candidate trajectories, a Safety Oracle stack that evaluates safety, and an Enforcement layer that applies risk policies at runtime. A governance MAS oversees monitoring, red‑teaming, verification, and versioned release management, enabling patch‑local fixes to safety failures without retraining the Proposer. The architecture is implementation‑agnostic and is demonstrated in two scenarios: a learned spatial Oracle and a clinical GenAI proxy. The authors provide open‑source code at https://github.com/decide-ugent/Alignment-Flywheel.

By Elias Malomgr\'e, Pieter Simoens