arXiv AI By Liyan Chen, Yael Tauman Kalai, Zoe Xi

How to Avoid Debate: Scalable AI Safety via Doubly-Efficient Interactive Proofs

Read the original on arXiv AI →

arXiv:2607. 03561v1 Announce Type: new Abstract: As AI models continue to develop powerful capabilities, it becomes critical that we are able to verify that their output is aligned with our intentions.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
2d ago

Can AI Oversight Be Zero Knowledge?

The paper investigates whether interactive arguments for oracle‑aided AI computations can be zero‑knowledge, meaning the verifier learns nothing beyond the correctness of the output. It proves that, in general, zero‑knowledge proofs for all oracle‑aided computations are impossible, even in the random oracle model, and this impossibility extends to debate protocols. However, if the oracle signs each answer with a cryptographic signature, then every oracle‑aided computation can be verified in zero‑knowledge with efficient provers and verifiers, assuming only collision‑resistant hash functions.

By Alessandro Chiesa, Ziyi Guan, Burcu Yildiz
arXiv AI
Jul 23

Avoiding Obfuscation with Prover-Estimator Debate

arXiv:2506. 13609v2 Announce Type: replace Abstract: Training powerful AI systems to exhibit desired behaviors hinges on the ability to provide accurate human supervision on increasingly complex tasks.

By Jonah Brown-Cohen, Geoffrey Irving, Georgios Piliouras, Lijie Chen, Jiawei Li, Zhiyang Xun