arXiv AI By Krishnapriya Vishnubhotla, Hillary Dawkins, Isar Nejadgholi, Svetlana Kiritchenko

Safety Measurements for Fine-tuned LLMs Should be Grounded in Capability

Read the original on arXiv AI →

arXiv:2606. 03648v1 Announce Type: cross Abstract: Adapting foundation large language models to a user's task or preferred style through fine-tuning can result in compromising the model's safety.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 19

Benchmarking the Benchmarks: Evaluating Automated Safety Benchmarks for Small Language Models

The paper investigates whether existing AI safety benchmarks, designed for large language models, are suitable for evaluating small language models (SLMs). By testing five benchmark suites on 26 open‑source SLMs with a unified scoring rubric, the authors find that ambiguous judgments dominate, especially for complex prompts and certain architectures. This ambiguity, linked to factors like lexical density and output perplexity, undermines the reliability of aggregate leaderboards and reveals a confound between model capability and perceived safety.

By Nyamtulla Shaik, Fengjun Li, Bo Luo
arXiv AI
Sep 25

Beyond Average Safety: Chance-Constrained LLM Fine-tuning

The paper introduces a chance-constrained approach to fine‑tune large language models (LLMs) that limits the proportion of safety examples whose performance degrades beyond a set threshold relative to a reference model. By replacing the discontinuous violation indicator with a differentiable majorization, the authors derive a tractable, conservative constraint and a closed‑form, constraint‑aware gradient update that focuses on examples near or above the degradation threshold. Experiments on harmful fine‑tuning across three tasks and models show that this tail‑aware method consistently outperforms existing safety‑preserving baselines, suggesting that safety preservation should be treated as a reliability‑constrained optimization problem rather than average‑risk regularization.

By Taha Entesari, Mahyar Fazlyab