arXiv Machine Learning By Aleksander Fafu{\l}a

Abliteration Is Not a Scalpel: Off-Target Effects of Refusal Removal on Decision Disposition Across Model Families

Read the original on arXiv Machine Learning →

arXiv:2607. 17427v1 Announce Type: new Abstract: Abliteration - deleting a model's refusal direction from its weights - is the standard recipe behind popular "uncensored" open-weight models.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.

arXiv AI
5d ago

Dead text or binding clause? Measuring and restoring constraint influence in black-box LLM dialogues

arXiv:2608. 12599v1 Announce Type: new Abstract: Multi-turn dialogues let users revoke constraints as easily as impose them, but revocation does not reliably take effect: models keep enacting withdrawn requirements (occasionally beneath comments asserting their removal), a failure we call \emph{behavioral relapse}, or revocation inertia.

By Haoyuan Zhu
arXiv AI
Jul 15

A Threshold Exceedance Framework for CBRN Uplift Evaluation in Frontier Language Models

arXiv:2607. 12200v1 Announce Type: new Abstract: As frontier language models advance, policymakers and model developers need methods for assessing whether model access materially increases a non-expert actor's ability to plan high-consequence Chemical, Biological, Radiological, or Nuclear (CBRN) misuse relative to public tools alone.

By Rahul Gupta, Abhinav Mohanty, Payal Motwani, Venkatesh Saligrama, Satyapriya Krishna, Connor Harris, Gary Anthony Ackerman, Brandon Behlendorf, Tom Hobson, Theodore Wilson, Spyros Matsoukas