arXiv AI

Palette: A Modular, Controllable, and Efficient Framework for On-demand Authorized Safety Alignment Relaxation in LLMs

arXiv AI
Sep 18

From Intent to Action: Benchmarking LLM Safety in Vehicle Voice Command Authorization

The paper introduces a 202-scenario benchmark to evaluate how large language models (LLMs) handle safety-critical authorization decisions for vehicle voice commands. It tests two local open-weight models and three API-based LLMs, finding alignment scores ranging from 40.1% to 89.1% and noting persistent false execution errors. The study concludes that structured LLM decisions alone are insufficient for safety, recommending an independent enforcement layer to verify tool permissions and vehicle-state constraints before any vehicle function is invoked.

By Diba Afroze, Xingli Zhang, Yazhou Tu, Xiali Hei
arXiv AI
Jun 6

Do Models Share Safety Representations? Cross-Model Steering for Safe Visual Generation

arXiv:2606. 05290v1 Announce Type: cross Abstract: Recent progress in generative modeling has made safety control a central challenge, yet existing approaches remain largely model-specific, requiring retraining or tailored interventions for each new architecture.

By Tobia Poppi, Silvia Cappelletti, Sara Sarto, Florian Schiffers, Garin Kessler, Marcella Cornia, Lorenzo Baraldi, Rita Cucchiara
arXiv Computation and Language
Aug 25

Safety-Aligned Weights Are Not Enough: Refusal-Teacher-Guided Finetuning Enhances Safety and Downstream Performance under Harmful Finetuning Attacks

arXiv:2506.07356v3 Announce Type: replace Abstract: While Finetuning-as-a-Service (FaaS) enables customization of Large Language Models (LLMs) using user data, this service is vulnerable to safety de...

By Seokil Ham, Yubin Choi, Yujin Yang, Seungju Cho, Younghun Kim, Changick Kim
arXiv Machine Learning
Sep 22

SafeTune: A Unified Faithful Library for Auditing and Repairing Safety Drift in Fine-Tuned LLMs

SafeTune is a source‑available library that consolidates four safety‑intervention paradigms—post‑hoc weight recovery, safety‑constrained fine‑tuning, gradient‑based unlearning, and inference‑time steering—into a single, configuration‑driven workflow. It offers shared interpretability, evaluation, and deployment tools, and its modular registry allows easy addition of new methods, benchmarks, judges, models, and fine‑tuning domains. The authors demonstrate SafeTune with controlled comparisons and case studies in finance and medical deployments, showing how it characterizes safety drift, evaluates interventions on refusal‑behavior and capability metrics, and supports calibrated or layered mitigation.

By Pratinav Seth, Saisab Sadhu, Anshul Kaushal, Vinay Kumar Sankarapu