← Back to all news
arXiv AI September 2, 2026 By Kuan-Lin Chu, Chung-En Sun, Tsui-Wei Weng

From Detection to Refusal: Safer LLMs via Circuit-Guided Weight Scaling

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

  • llms
  • benchmarks
  • safety

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv AI
Aug 24

RefusalGuard: Geometry-Preserving Fine-Tuning for Safety in LLMs

arXiv:2605.01913v2 Announce Type: replace-cross Abstract: Fine-tuning safety-aligned language models for downstream tasks often leads to substantial degradation of refusal behavior, making models vul...

By Sadia Asif, Mohammad Mohammadi Amiri
llmsfine-tuningbenchmarkssafety
More like this →
arXiv Machine Learning
Jul 13

Optimizing Against Safety Representations: Activation-Guided Adversarial Suffixes and the Geometry of Refusal

arXiv:2607. 08883v1 Announce Type: new Abstract: Behavioral alignment in large language models often masks fragile internal safety representations.

By Ege \c{C}akar, Hannah Guan, Kayden Kehe
llmssafety
More like this →
arXiv AI
Aug 17

Tripwire: Triggering Aligned Refusal via Statistically Certified Safety Neurons

arXiv:2608. 14392v1 Announce Type: new Abstract: Neuron- and path-level interventions offer the finest-grained route to defending large language models (LLMs) against jailbreak attacks, yet existing methods fall short of this promise, i.

By Wei Zhao, Zhe Li, Peixin Zhang, Jun Sun
llmssafety
More like this →
arXiv AI
Jul 23

Defense Against LLM Backdoors using Critical Neuron Isolation Pruning

arXiv:2607. 19894v1 Announce Type: cross Abstract: Large language models (LLMs) are vulnerable to backdoor attacks, where hidden triggers induce malicious outputs.

By Yuxi Li, Zhibo Zhang, Kailong Wang, Xingshuo Han, Ling Shi, Haoyu Wang
llmsfine-tuningefficiencymultimodalbenchmarks
More like this →
Hugging Face Trending Papers
Jul 22

Defense Against LLM Backdoors using Critical Neuron Isolation Pruning

Large language models (LLMs) are vulnerable to backdoor attacks, where hidden triggers induce malicious outputs. Existing defenses generally fall into inference-time detection or training-time mitigation, but face two key limitations.

llmsfine-tuningefficiencymultimodalbenchmarks
More like this →
arXiv AI
Jun 6

Safety Paradox: How Enhanced Safety Awareness Leaves LLMs Vulnerable to Posterior Attack

arXiv:2606. 05614v1 Announce Type: new Abstract: Large language models (LLMs) are rigorously aligned to refuse harmful requests, a process that inherently cultivates a latent capacity to evaluate and recognize unsafe content.

By Long P. Hoang, Hai V. Le, Shaoyang Xu, Wei Lu, Wenxuan Zhang
llmsreinforcement-learningsafety
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea