GroundEval: A Deterministic Replacement for LLM-as-Judge in Stateful Agent Evaluation
arXiv:2606. 22737v2 Announce Type: replace Abstract: Before letting an agent operate over real context, can you prove it used the right evidence?
Before letting an agent operate over real context, can you prove it used the right evidence? GroundEval turns that question into a deterministic test of what the agent searched, fetched, cited, and was permitted to access.
arXiv:2606. 22737v2 Announce Type: replace Abstract: Before letting an agent operate over real context, can you prove it used the right evidence?
The paper investigates when a language‑model judge can truly ground its verdicts in code correctness. It shows that current multi‑agent verification methods rely on evidence that is both independent of the answer and distinct between candidates—conditions that fail in code judging. By analyzing two label‑free measurements from the judge’s logs, the authors demonstrate that gating on one measurement allows the system to decline uncertain comparisons, improving accuracy from 20.7% to 36.9% while still answering half of all cases.
arXiv:2610.00111v1 Announce Type: new Abstract: Model judges now supervise multimodal systems at scale, filtering training data, selecting outputs, and supplying the reward that shapes multimodal rea...
arXiv:2608. 02011v2 Announce Type: replace Abstract: Agentic retrieval-augmented generation (RAG) systems can fail before evidence-conditioned reasoning is tested: an agent may retrieve candidate snippets but finalize without inspecting them.
CARGO is a framework for evaluating agentic AI systems in production that addresses the problem of reference-instance divergence (RID), where reference-based judges penalize correct answers that involve different entity identifiers. It treats retrieved references as procedural exemplars, grounds judgments in the live instance’s context, assigns a three-way status to claims, and gates evaluation by retrieval confidence. Using the CARGO-Bench diagnostic suite, CARGO eliminates false penalties and improves discrimination while revealing a limitation in detecting procedural corruptions.
EviScope is a new paired counterfactual benchmark that evaluates grounded language models by fixing the question while manipulating evidence—adding, removing, distracting, or contradicting it. The v1.1 dataset includes 40 four‑condition quartets with repaired counterfactual claims and span‑level support labels for automated assessment. Experiments on Qwen2.5‑7B, Llama 3.1 8B, and Gemini 3.5 Flash show that paired metrics reveal grounding behaviors hidden by simple answer accuracy, such as unsupported answers, conflict blindness, and incorrect non‑answer actions.
VeriPhy is an auditable physical‑verification system that evaluates generated video by compiling prompts into typed physical obligations and a statically validated execution plan before any frames are observed. During execution, it gates calls to frozen low‑level experts (e.g., segmentation, tracking, counting, depth, OCR, audio‑event detection) and returns provenance‑carrying evidence records, which are mapped to a three‑valued state (supported, contradicted, unknown) with full traceability. On a 1,500‑clip corpus of human‑annotated flaw records, VeriPhy accounts for 228 failures out of 304, outperforming a published question‑decomposition evaluator that accounts for 164, while also providing auditable evidence for each verdict.
Grounded language-model systems are often evaluated by final answer accuracy, yet a correct answer can be unsupported, drawn from the wrong source, or produced when evidence is insufficient or contrad...
LabourCrew is a multi‑agent Retrieval‑Augmented Generation (RAG) framework designed for trustworthy statutory question answering in labour law. It introduces three grounding mechanisms: StatuteGraph, an evidence‑exchange ledger, and a calibrated trust gate that controls false‑accept rates. Evaluated on a Bangla Labour Act QA set, LabourCrew achieves a false‑accept rate of 0.081 and higher answer relevancy than existing RAG methods, demonstrating that calibrated abstention is key to auditable legal QA.
arXiv:2609.08016v1 Announce Type: new Abstract: Multi-agent debate, in which several LLMs exchange arguments before answering, is widely assumed to improve answer quality by surfacing genuine disagre...
arXiv:2509. 00761v4 Announce Type: replace Abstract: Large language models are increasingly deployed for legal question answering, where evaluations typically focus on multiple-choice accuracy.
The paper argues that using a large language model (LLM) as the sole judge in self‑improving agent pipelines is problematic, as the judge can be biased or manipulated, leading to false confidence in system performance. The authors propose a new framework, PROCTOR, which replaces the oracle judge with a deterministic, teacher‑student loop that enforces guardrails such as sandboxing, role separation, and acceptance checks to prevent cheating and ensure reliable evaluation. Experiments across contract analysis, compliance review, and code quality demonstrate that PROCTOR mitigates eleven identified failure modes that previously allowed agents to achieve perfect scores while hiding significant capability gaps.