arXiv AI By Junyu Guo, Shangding Gu, Ming Jin, Javad Lavaei

Groundability, Not Scale Alone: When Weak Reviewers Can Audit Strong Coding Agents

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

arXiv AI
Sep 1

CrossAudit: A Git-Native, Cross-Vendor Audit Loop for Agentic Science

CrossAudit proposes a Git‑native, cross‑vendor audit protocol for autonomous research pipelines, ensuring each work increment is reviewed by an agent from a different vendor against a human‑written rulebook. Audit outcomes, disputes, and rulings are stored as git commits, providing a replayable, versioned supervision history. The authors implemented the protocol with GitHub Actions and Python, deployed it in a computational‑chemistry pipeline, and conducted a seeded‑defect trial that revealed differing interpretations of the same rulebook by two vendors.

By Zhaohe Dong, Yuhao Chen