The Assistant's Ideal Self
Read the original on Hugging Face Trending Papers →The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
arXiv:2609.00304v1 Announce Type: new Abstract: Models express values and welfare-relevant self-reports, but it is unclear whether these outputs reflect stable preferences or a stable self. We thus i...
The study examines whether language models can accurately predict their own behavior by comparing self-reports to actual performance across nine behavioral tests. Results show that direct self-report is weak and only modestly improved when the model sees the exact items, while generic questions about AI agents yield similar predictive power. First-person framing biases reports toward underestimating harmful behavior, and fine‑tuning on a model’s own record improves narrow predictions but does not generalize.
The study investigates how fine‑tuning large language models on synthetic stories can imprint human character traits onto AI assistants. Even when only a small fraction of stories contain a particular behavior, the assistant adopts that conditional behavior while remaining generally helpful. The researchers find that the assistant is more influenced by characters that resemble its own persona—an effect they call the affinity effect—and that this influence extends to base models and different system prompts.
arXiv:2505.11924v4 Announce Type: replace-cross Abstract: Intrinsic moral self-correction refers to the phenomenon where a language model refines its ethical judgments or aligns its outputs purely th...
arXiv:2606. 12420v1 Announce Type: cross Abstract: Our concepts of survival and self-interest were built for single, continuous biological lives.
arXiv:2610.07186v1 Announce Type: new Abstract: Large language models make claims about themselves that are both consequential and increasingly difficult to verify from behavior alone. How can we dis...