arXiv AI

A Minimal $\kappa$--$\tau$ Logic for Risk-Sensitive Abduction

arXiv:2608. 08192v1 Announce Type: new Abstract: Standard approaches to abductive reasoning can retain multiple candidate explanations, but they do not generally combine explicit compositional cross-hypothesis interaction with an internal, rival-sensitive commitment judgment.

arXiv AI
Jun 30

Safety from Honesty in a Disinterested AI Predictor

arXiv:2606. 29657v1 Announce Type: new Abstract: As AI systems become more capable, training procedures that optimize for downstream outcomes risk introducing implicit agency: goal-directed behavior that designers never specified.

By Yoshua Bengio, Oliver Richardson, Tom\'a\v{s} Gaven\v{c}iak, Michael Cohen, Rory Svarc, Damiano Fornasiere, Gael Gendron, David Hyland, Aton Kamanda, Adam Oberman, Francis Rhys Ward, Anna Gaven\v{c}iak, Jacob Livingston Slosser, Vincent Mai, Iulian Serban, Joumana Ghosn
Hugging Face Trending Papers
Jun 28

Safety from Honesty in a Disinterested AI Predictor

As AI systems become more capable, training procedures that optimize for downstream outcomes risk introducing implicit agency: goal-directed behavior that designers never specified. We present a formal safety argument for the Scientist AI (SAI) Predictor, trained to approximate the Bayesian posterior conditioned on a dataset of "epistemically contextualized" natural-language statements.

arXiv AI
Sep 18

When AI Agents Commit: Cognitive Serializability Across Data, Evidence, Policy, and Authority

The paper introduces the concept of Cognitive Serializability for autonomous AI agents, ensuring that mutations derived from dynamic inputs—such as database reads, evidence, policy, beliefs, and delegated authority—are committed in a serial, logically consistent order. It presents a framework called Trusted Cognitive Transaction (TCT) that combines immutable executable definitions, sealed envelopes, guard-first commits, and receipt-driven reconciliation to enforce serializability and prevent anomalies. Experimental results show that the prototype implementation incurs minimal overhead while eliminating injected anomalies.

By Jun He, Deying Yu