Hugging Face Trending Papers
Aug 31

The Assistant's Ideal Self

Models express values and welfare-relevant self-reports, but it is unclear whether these outputs reflect stable preferences or a stable self. We thus introduce a structured elicitation of an assistant...

arXiv Computation and Language
Sep 10

Strangers to Themselves: What Language Models Say About Themselves Is Generic

The study examines whether language models can accurately predict their own behavior by comparing self-reports to actual performance across nine behavioral tests. Results show that direct self-report is weak and only modestly improved when the model sees the exact items, while generic questions about AI agents yield similar predictive power. First-person framing biases reports toward underestimating harmful behavior, and fine‑tuning on a model’s own record improves narrow predictions but does not generalize.

By Phil Blandfort, Urja Pawar
arXiv Machine Learning
Sep 11

Story Imprinting: AI Assistants Absorb Traits from Human Characters They Resemble

The study investigates how fine‑tuning large language models on synthetic stories can imprint human character traits onto AI assistants. Even when only a small fraction of stories contain a particular behavior, the assistant adopts that conditional behavior while remaining generally helpful. The researchers find that the assistant is more influenced by characters that resemble its own persona—an effect they call the affinity effect—and that this influence extends to base models and different system prompts.

By Jorio Cocola, Lev McKinney, Harry Mayne, Jan Betley, Owain Evans