← Back to all news
Hugging Face Trending Papers September 29, 2026

Safer Content or Firmer Refusals? A Hybrid Perturbation Defense for Alignment under Harmful Fine-tuning

Read the original on Hugging Face Trending Papers →

The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.

  • llms
  • rag
  • fine-tuning
  • safety

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv AI
2d ago

Safer Content or Firmer Refusals? A Hybrid Perturbation Defense for Alignment under Harmful Fine-tuning

arXiv:2609.36862v1 Announce Type: cross Abstract: Fine-tuning-as-a-service lets users adapt a safety-aligned language model to their own data, but it also creates a harmful fine-tuning attack surface...

By Muhammad Zeeshan Akram, Mufid Kamel Marican, Anvesh Reddy Yenugu, Ali Zain Kaimkhani, Minghong Fang
llmsragfine-tuningsafety
More like this →
arXiv AI
Aug 13

Localizing Safety Alignment: MLP Layers and Mid-Network Blocks Encode Refusal Behavior in Large Language Models

arXiv:2608. 11583v1 Announce Type: new Abstract: Safety alignment in large language models is often treated as a distributed property of the entire network, yet its practical brittleness suggests that refusal behavior may be concentrated in a smaller set of parameters.

By Mingyu Zong, Sampad Mohanty, Bhaskar Krishnamachari
llmsbenchmarkssafety
More like this →
arXiv Machine Learning
Jun 4

Few Tokens, Big Leverage: Preserving Safety Alignment by Constraining Safety Tokens during Fine-tuning

arXiv:2603. 07445v2 Announce Type: replace-cross Abstract: Large language models (LLMs) often require fine-tuning (FT) to perform well on downstream tasks, but FT can induce safety-alignment drift even when the training dataset contains only benign data.

By Guoli Wang, Haonan Shi, Tu Ouyang, An Wang
llmsfine-tuningsafety
More like this →
arXiv AI
Jul 2

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment

arXiv:2607. 00572v1 Announce Type: new Abstract: Understanding how aligned LLMs internally represent safety is critical for diagnosing alignment vulnerabilities, as it explains why jailbreaks succeed and informs the design of robust alignment strategies.

By Shei Pern Chua, Fangzhao Wu
llmsfine-tuningsafety
More like this →
arXiv AI
Jun 18

Beyond Safe Data: Pretraining-Stage Alignment with Regular Safety Reflection

arXiv:2606. 19168v1 Announce Type: new Abstract: To achieve deeper safety alignment for large language models (LLMs), recent efforts have studied how to push safety interventions earlier into the pretraining stage, primarily by filtering unsafe data or rewriting it into safer forms.

By Jinhan Li, Kexian Tang, Yihan Xu, Zhuorui Ye, Kaifeng Lyu
llmsfine-tuningsafety
More like this →
arXiv AI
Jul 23

Matching Ranks Over Probability Yields Truly Deep Safety Alignment

arXiv:2512. 05518v2 Announce Type: replace-cross Abstract: Open-source Large Language Models (LLMs) play a critical role in the democratization of AI, yet their "open" nature introduces more avenues for malicious actors to misuse them for harmful purposes.

By Jason Vega, Gagandeep Singh
llmsfine-tuningsafety
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea