arXiv AI

Low-Agreeableness Persona Conditioning for Safe LLM Fine-Tuning

arXiv:2606. 27709v1 Announce Type: cross Abstract: Recent work has shown that fine-tuning large language models (LLMs) for social warmth degrades factual reliability and increases sycophancy.

arXiv AI
Jul 3

Psychological Steering in LLMs: An Evaluation of Effectiveness and Trustworthiness

arXiv:2510. 04484v2 Announce Type: replace-cross Abstract: The ability to control LLMs' emulated emotional states and personality traits is an essential step in enabling rich, human-centered interactions in socially interactive settings.

By Amin Banayeeanzade, Ala N. Tak, Fatemeh Bahrani, Anahita Bolourani, Leonardo Blas, Emilio Ferrara, Jonathan Gratch, Sai Praneeth Karimireddy
arXiv AI
Sep 24

Beyond Unsafe Detection: Counterfactually Anchored Evidence Attribution for Multi-Turn LLM Safety Failures

The paper introduces a counterfactually anchored evidence attribution approach for multi‑turn large language model safety failures. It presents a new dataset of 1,762 conversations, including adversarial, benign twins, and high‑risk vocabulary variants, and trains a lightweight hierarchical model that accurately predicts safety violations and attributes them to specific user turns and token spans. The model achieves high detection performance (F1 = 0.988) and significantly reduces adversarial confidence when top‑attributed tokens are removed, while maintaining low false‑positive rates on benign conversations.

By Srinivasan Subramanian, Kazi Aminul Islam, Md. Abdullah Al Hafiz Khan
arXiv Computation and Language
3d ago

Evaluating Language Model Safety Across Long Adversarial Conversations

The paper investigates how conversational safety in language models degrades over extended, adversarial interactions. By testing three instruction‑tuned models with persistent adversarial users across up to 101 turns, the study finds that safe‑response rates drop sharply from 85–100% at the first turn to 15–44% by the end. This demonstrates that strong single‑turn safety does not guarantee continued safety in long conversations.

By Parisa Salmani, Peter R. Lewis
arXiv Computation and Language
3d ago

LLM Persona Unlearning

arXiv:2609.39882v1 Announce Type: new Abstract: Pre-training equips large language models (LLMs) with a broad repertoire of behavioral patterns associated with roles, styles, values, and goals. Post-...

By Kemou Li, Zhuan Shi, Qizhou Wang, Fengpeng Li, Negar Rostamzadeh, Golnoosh Farnadi, Jiantao Zhou
arXiv AI
Jul 10

Persona Cartography: Charting Language Model Personality Traits in Weight Space

arXiv:2607. 07916v1 Announce Type: new Abstract: Large language models exhibit recurring behavioural patterns -- personas -- that shape generalisation and safety, but we lack reliable tools for decomposing, measuring, and controlling them.

By Luke Baines, Anton Gonzalvez Hawthorne, Mariia Koroliuk, Irakli Shalibashvili, Cl\'ement Dumas, Konstantinos Voudouris, David Demitri Africa
arXiv AI
Sep 7

Refuse without Refusal: A Structural Analysis of Safety-Tuning Responses for Reducing False Refusals in Language Models

The paper investigates how large language models balance helpfulness and safety by refusing harmful queries while responding to benign ones. It decomposes safety-tuning responses into a boilerplate refusal statement and a rationale, finding that the statement causes false refusals by relying on superficial cues. Training on rationales alone reduces false refusals without compromising safety performance, suggesting that fine‑grained safety supervision is essential for better alignment.

By Minji Kim, Hyounghun Kim