arXiv AI By Benjamin Sturgeon, David Africa, Sid Black

When Role-playing, Do Models Believe What They Say?

Read the original on arXiv AI →

arXiv:2606. 11502v3 Announce Type: replace-cross Abstract: Language models can state that "the Earth orbits the Sun" and, when role-playing Aristotle, assert the opposite.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Jul 8

Dissociating the Internal Representations of Sycophancy in LLMs

Large Language Models (LLMs) frequently exhibit sycophancy, where they agree with a user's statement even when incorrect. While sycophancy is often treated as a single defined behavior, it can manifest in substantially distinct ways and circumstances, raising the question of whether this multi-faceted nature is reflected in its internal mechanisms.

arXiv Machine Learning
Sep 11

Story Imprinting: AI Assistants Absorb Traits from Human Characters They Resemble

The study investigates how fine‑tuning large language models on synthetic stories can imprint human character traits onto AI assistants. Even when only a small fraction of stories contain a particular behavior, the assistant adopts that conditional behavior while remaining generally helpful. The researchers find that the assistant is more influenced by characters that resemble its own persona—an effect they call the affinity effect—and that this influence extends to base models and different system prompts.

By Jorio Cocola, Lev McKinney, Harry Mayne, Jan Betley, Owain Evans