arXiv Computation and Language
Aug 25

Expectations and Practices around AI Disclosure in CS Research

The paper examines AI disclosure policies in top computer science venues, finding them to be highly under‑specified. A survey of 109 researchers shows that disclosures are deemed most necessary for research design tasks and when human involvement is low, and it compiles researchers’ expectations for disclosure content. Analysis of 13,867 disclosure statements from EMNLP 2025 and ICLR 2026 reveals a significant mismatch between these expectations and actual practice, such as frequent disclosure of writing assistance despite it being considered less necessary.

By Arati Mohapatra, Danish Pruthi
arXiv AI
Jun 30

Safety from Honesty in a Disinterested AI Predictor

arXiv:2606. 29657v1 Announce Type: new Abstract: As AI systems become more capable, training procedures that optimize for downstream outcomes risk introducing implicit agency: goal-directed behavior that designers never specified.

By Yoshua Bengio, Oliver Richardson, Tom\'a\v{s} Gaven\v{c}iak, Michael Cohen, Rory Svarc, Damiano Fornasiere, Gael Gendron, David Hyland, Aton Kamanda, Adam Oberman, Francis Rhys Ward, Anna Gaven\v{c}iak, Jacob Livingston Slosser, Vincent Mai, Iulian Serban, Joumana Ghosn
Hugging Face Trending Papers
Jun 28

Safety from Honesty in a Disinterested AI Predictor

As AI systems become more capable, training procedures that optimize for downstream outcomes risk introducing implicit agency: goal-directed behavior that designers never specified. We present a formal safety argument for the Scientist AI (SAI) Predictor, trained to approximate the Bayesian posterior conditioned on a dataset of "epistemically contextualized" natural-language statements.