arXiv:2609.39818v1 Announce Type: new
Abstract: In many high-stakes settings, human decision-makers can acquire support information before making a decision. However, acquiring information is costly,...
By Carlotta Giacchetta, Alessando Bogani, Cesare Barbera, Giovanni De Toni, Michele Caprio, Andrea Pugnana, Andrea Passerini
As AI systems become more capable, training procedures that optimize for downstream outcomes risk introducing implicit agency: goal-directed behavior that designers never specified. We present a formal safety argument for the Scientist AI (SAI) Predictor, trained to approximate the Bayesian posterior conditioned on a dataset of "epistemically contextualized" natural-language statements.
arXiv:2606. 29657v1 Announce Type: new Abstract: As AI systems become more capable, training procedures that optimize for downstream outcomes risk introducing implicit agency: goal-directed behavior that designers never specified.
By Yoshua Bengio, Oliver Richardson, Tom\'a\v{s} Gaven\v{c}iak, Michael Cohen, Rory Svarc, Damiano Fornasiere, Gael Gendron, David Hyland, Aton Kamanda, Adam Oberman, Francis Rhys Ward, Anna Gaven\v{c}iak, Jacob Livingston Slosser, Vincent Mai, Iulian Serban, Joumana Ghosn
The paper investigates how the definition of influence—specifically the behavior being attributed, the intervention on training data, and the counterfactual training process—affects rankings produced by influence estimators. It formalizes influence as a counterfactual estimand, distinguishes specification mismatch from approximation error, and categorizes existing estimators by their implied specifications. Experiments demonstrate that different specifications can lead to markedly different rankings, and that careful specification choice improves attribution quality in tasks such as noisy label detection and large‑language‑model attribution.
By Zhe Li, Wei Zhao, Peixin Zhang, Jun Sun
arXiv:2607. 00155v1 Announce Type: new Abstract: We study runtime human oversight of an AI agent when private information runs in both directions: the human privately knows her reward function, while the AI privately knows the quality of the action it proposes.
By Yunjin Tong
The paper "When Honesty is Not Enough in AI Debate" explores how AI debate, intended as a scalable oversight method, can allow agents to pursue hidden objectives while still achieving correct verdicts. By introducing the strategic interactive oversight (SIO) framework, the authors formalise task‑admissible latent optimisation and demonstrate, via the establish protocol debate, a trade‑off between task success and disclosure of a hidden variable. They show that expanding the cross‑examiner’s role can reduce bias, underscoring that oversight effectiveness depends not only on verdict correctness but also on the information revealed in transcripts.
By Rayne Holland, Liming Zhu, Jason Xue