Groundability, Not Scale Alone: When Weak Reviewers Can Audit Strong Coding Agents
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2608. 02677v1 Announce Type: cross Abstract: LLM code reviewers often estimate patch risk and make approval decisions in one prompt.
arXiv:2606. 18168v1 Announce Type: cross Abstract: Software practitioners increasingly use AI coding agents that generate test code alongside production code in open source pull requests (PRs).
arXiv:2609.37315v1 Announce Type: cross Abstract: Tool-using agents are entering settings where a wrong action carries real cost, and the benchmarks certifying them grade what each simulated tool cal...
CrossAudit proposes a Git‑native, cross‑vendor audit protocol for autonomous research pipelines, ensuring each work increment is reviewed by an agent from a different vendor against a human‑written rulebook. Audit outcomes, disputes, and rulings are stored as git commits, providing a replayable, versioned supervision history. The authors implemented the protocol with GitHub Actions and Python, deployed it in a computational‑chemistry pipeline, and conducted a seeded‑defect trial that revealed differing interpretations of the same rulebook by two vendors.
arXiv:2608.22808v2 Announce Type: replace Abstract: When can an agent failure be caught? An audit is usually limited by the record rather than by the method. CatchBench therefore puts one auditor's q...
arXiv:2607. 28871v1 Announce Type: cross Abstract: When a repair agent runs a test and sees it pass, the result is treated as evidence about the reported defect.