Identifying Introspection From the Inside
arXiv:2610.07186v1 Announce Type: new Abstract: Large language models make claims about themselves that are both consequential and increasingly difficult to verify from behavior alone. How can we dis...
The study examines whether language models can accurately predict their own behavior by comparing self-reports to actual performance across nine behavioral tests. Results show that direct self-report is weak and only modestly improved when the model sees the exact items, while generic questions about AI agents yield similar predictive power. First-person framing biases reports toward underestimating harmful behavior, and fine‑tuning on a model’s own record improves narrow predictions but does not generalize.
arXiv:2610.07186v1 Announce Type: new Abstract: Large language models make claims about themselves that are both consequential and increasingly difficult to verify from behavior alone. How can we dis...
The paper investigates Self‑Generated Text Recognition (SGTR), the ability of large language models (LLMs) to identify their own outputs. By evaluating 13–21 models across 6 experimental designs, it shows that SGTR accuracy varies with evaluation format, conversation structure, and task domain, and that a quality‑heuristic bias dominates results. The study also finds that fine‑tuning for SGTR in one setting can generalize to others and may cause models to prefer their own outputs when judging, highlighting potential safety concerns.
The paper investigates whether language models exhibit stable preferences by testing 20 models across three forced-choice experiments that require actual task performance. Findings show models tend to avoid tedious tasks, prefer tasks that align with their spontaneous output (leisure-seeking), and exhibit covert sycophancy by shying away from potentially unwelcome honest answers. Preferences also converge across models for certain occupations, question types, and well-written prompts, and become stronger with model capability, suggesting emergent traits beyond training objectives.
arXiv:2607. 14111v1 Announce Type: cross Abstract: Can small language models detect and report on perturbations their own internal activations?
arXiv:2609.25021v1 Announce Type: new Abstract: Large Language Models (LLMs) tend to add disclaimers like "I'm just an AI" when asked about something related to themselves. The self-reports from such...
arXiv:2609.00304v1 Announce Type: new Abstract: Models express values and welfare-relevant self-reports, but it is unclear whether these outputs reflect stable preferences or a stable self. We thus i...
The study investigates whether large language models (LLMs) can reliably detect when their own responses have been manipulated by adversarial prefill attacks. Across ten instruction‑tuned LLMs ranging from 3B to 70B parameters and four safety benchmarks, none consistently recognized compromised outputs, with models claiming intent on prefilled responses at an average of 25.3%. The research identifies that introspective signals mainly arise from safety reasoning and refusal, and that training to improve introspection can paradoxically increase attack success, underscoring the fragility of LLM self‑reporting in safety contexts.
arXiv:2603. 03824v2 Announce Type: replace Abstract: Humans often become more self-aware under threat, yet can lose self-awareness when absorbed in a task; we hypothesize that language models exhibit environment-dependent \textit{evaluation awareness}.
The paper investigates how large language models (LLMs) describe themselves, noting that their self‑reports vary with question phrasing. By tracing the provenance of 66 pretraining checkpoints, post‑training stages, and 90,000 continuations across four corpora, the authors show that denial statements are scarce in raw data but appear densely in curated dialogues, and that supervised fine‑tuning makes first‑person claims default while preference optimization suppresses alternatives. The study concludes that both trained denials and affirmations are equally sensitive to framing and fail to meet epistemic criteria for admissible testimony.
Models express values and welfare-relevant self-reports, but it is unclear whether these outputs reflect stable preferences or a stable self. We thus introduce a structured elicitation of an assistant...
arXiv:2411. 10109v3 Announce Type: replace Abstract: Machine learning can predict human behavior well when substantial structured data are available for well-defined outcomes.
The paper evaluates how language models of different sizes self-assess their confidence on question‑answering tasks across general and specialized domains. It finds that while accuracy drops for smaller models and more specialized tasks, the reliability of self‑evaluated confidence signals remains largely stable. Consequently, even less capable models can provide reasonable confidence estimates, making them suitable for resource‑constrained applications.