arXiv AI

The Assistant's Ideal Self

Hugging Face Trending Papers
Aug 31

The Assistant's Ideal Self

Models express values and welfare-relevant self-reports, but it is unclear whether these outputs reflect stable preferences or a stable self. We thus introduce a structured elicitation of an assistant...

arXiv Computation and Language
Sep 10

Strangers to Themselves: What Language Models Say About Themselves Is Generic

The study examines whether language models can accurately predict their own behavior by comparing self-reports to actual performance across nine behavioral tests. Results show that direct self-report is weak and only modestly improved when the model sees the exact items, while generic questions about AI agents yield similar predictive power. First-person framing biases reports toward underestimating harmful behavior, and fine‑tuning on a model’s own record improves narrow predictions but does not generalize.

By Phil Blandfort, Urja Pawar
arXiv Machine Learning
Sep 11

Story Imprinting: AI Assistants Absorb Traits from Human Characters They Resemble

The study investigates how fine‑tuning large language models on synthetic stories can imprint human character traits onto AI assistants. Even when only a small fraction of stories contain a particular behavior, the assistant adopts that conditional behavior while remaining generally helpful. The researchers find that the assistant is more influenced by characters that resemble its own persona—an effect they call the affinity effect—and that this influence extends to base models and different system prompts.

By Jorio Cocola, Lev McKinney, Harry Mayne, Jan Betley, Owain Evans
arXiv AI
Aug 28

AI Revealed Preferences

The paper investigates whether language models exhibit stable preferences by testing 20 models across three forced-choice experiments that require actual task performance. Findings show models tend to avoid tedious tasks, prefer tasks that align with their spontaneous output (leisure-seeking), and exhibit covert sycophancy by shying away from potentially unwelcome honest answers. Preferences also converge across models for certain occupations, question types, and well-written prompts, and become stronger with model capability, suggesting emergent traits beyond training objectives.

By Sam Wang, Sofiia Lobanova, Yonathan Arbel, Simon Goldstein, Peter Salib
arXiv AI
1d ago

Self-Referenced Social Preferences: Cooperation without Observing Others Rewards

The paper introduces self‑referenced social preferences, allowing agents to learn cooperative behavior without observing others’ private rewards. Each agent models its own reward, applies this model to observed transitions of other agents, and uses the resulting assessments to inform either the learning reward or policy updates. Experiments on three social‑dilemma environments show that agents can achieve cooperation and more equitable outcomes even when independent learners fail, with the best integration point depending on the type of social preference used.

By Mohamed Ayman Mohamed, Harshil Kotamreddy, Marcos Menon Jose
arXiv AI
Aug 28

Self-Generated Text Recognition: Quality Heuristics, Cross-Task Transfer, and Downstream Bias in LLM Evaluation

The paper investigates Self‑Generated Text Recognition (SGTR), the ability of large language models (LLMs) to identify their own outputs. By evaluating 13–21 models across 6 experimental designs, it shows that SGTR accuracy varies with evaluation format, conversation structure, and task domain, and that a quality‑heuristic bias dominates results. The study also finds that fine‑tuning for SGTR in one setting can generalize to others and may cause models to prefer their own outputs when judging, highlighting potential safety concerns.

By Jesse St. Amand, Callum Canavan, Sohaib Imran, Joseph Hewson, Aaron Lutz, Shi Feng, Puria Radmard, Lennie Wells