Measuring the Assistant's Harmlessness Preferences on the User Turn
Read the original on arXiv Computation and Language →The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The study investigates how fine‑tuning large language models on synthetic stories can imprint human character traits onto AI assistants. Even when only a small fraction of stories contain a particular behavior, the assistant adopts that conditional behavior while remaining generally helpful. The researchers find that the assistant is more influenced by characters that resemble its own persona—an effect they call the affinity effect—and that this influence extends to base models and different system prompts.
arXiv:2607. 13162v1 Announce Type: cross Abstract: What a language model will and will not do is largely set during post-training, but which behaviors it expresses, hides, or resists is not revealed by prompting alone.
arXiv:2606. 12747v1 Announce Type: new Abstract: Safety-relevant studies of language models, including alignment and jailbreaking evaluations and AI control protocols, often rely on prefilling model outputs.
arXiv:2608. 13482v1 Announce Type: cross Abstract: As language-model-based AI is increasingly deployed in autonomous settings, aligning its goals and values with those of humans becomes critical.
arXiv:2606. 27709v1 Announce Type: cross Abstract: Recent work has shown that fine-tuning large language models (LLMs) for social warmth degrades factual reliability and increases sycophancy.
The paper introduces the Atomic User Model (AUM), a structured representation of a user’s personality that separates a stable identity nucleus from four interpretable shells—psychological, cognitive & experiential, behavioural, and social—along with cross-shell entries for conflict and authenticity. It proposes using AUM as a retrieval index rather than a prompt prefix, enabling a task‑specific, budgeted retrieval of relevant fields at generation time. Experiments with simulated participants show that retrieving eight AUM fields improves style fidelity, preference accuracy, and user voice identification compared to flat preference notes, especially benefiting users whose default assistant performs poorly.