arXiv AI
2d ago

Who Put the I in AI? Provenance and the Admissibility of Machine Self-Report

The paper investigates how large language models (LLMs) describe themselves, noting that their self‑reports vary with question phrasing. By tracing the provenance of 66 pretraining checkpoints, post‑training stages, and 90,000 continuations across four corpora, the authors show that denial statements are scarce in raw data but appear densely in curated dialogues, and that supervised fine‑tuning makes first‑person claims default while preference optimization suppresses alternatives. The study concludes that both trained denials and affirmations are equally sensitive to framing and fail to meet epistemic criteria for admissible testimony.

By Kristina \v{S}ekrst
arXiv AI
Aug 26

Can Large Language Models Still Explain Themselves? Investigating the Impact of Quantization on Self-Explanations

The paper investigates how quantization affects large language models’ self‑explanations, examining natural language explanations and counterfactual examples across three quantization techniques and bit widths. Results show moderate declines in explanation quality (up to 4.4%) and faithfulness (up to 3.9%), with user studies indicating up to an 8.5% drop in coherence and trustworthiness. Larger models are less resilient in quality but remain more faithful, and no single quantization method consistently outperforms others across accuracy, quality, and faithfulness.

By Qianli Wang, Nils Feldhus, Pepa Atanasova, Fedor Splitt, Simon Ostermann, Sebastian M\"oller, Vera Schmitt
arXiv AI
Aug 28

Self-Generated Text Recognition: Quality Heuristics, Cross-Task Transfer, and Downstream Bias in LLM Evaluation

The paper investigates Self‑Generated Text Recognition (SGTR), the ability of large language models (LLMs) to identify their own outputs. By evaluating 13–21 models across 6 experimental designs, it shows that SGTR accuracy varies with evaluation format, conversation structure, and task domain, and that a quality‑heuristic bias dominates results. The study also finds that fine‑tuning for SGTR in one setting can generalize to others and may cause models to prefer their own outputs when judging, highlighting potential safety concerns.

By Jesse St. Amand, Callum Canavan, Sohaib Imran, Joseph Hewson, Aaron Lutz, Shi Feng, Puria Radmard, Lennie Wells