Before letting an agent operate over real context, can you prove it used the right evidence? GroundEval turns that question into a deterministic test of what the agent searched, fetched, cited, and was permitted to access.
CARGO is a framework for evaluating agentic AI systems in production that addresses the problem of reference-instance divergence (RID), where reference-based judges penalize correct answers that involve different entity identifiers. It treats retrieved references as procedural exemplars, grounds judgments in the live instance’s context, assigns a three-way status to claims, and gates evaluation by retrieval confidence. Using the CARGO-Bench diagnostic suite, CARGO eliminates false penalties and improves discrimination while revealing a limitation in detecting procedural corruptions.
By Mukul Chhabra, Shail Patel, Luigi Medrano
arXiv:2608. 02011v2 Announce Type: replace Abstract: Agentic retrieval-augmented generation (RAG) systems can fail before evidence-conditioned reasoning is tested: an agent may retrieve candidate snippets but finalize without inspecting them.
By Daeyoung Roh, Donghee Han
VeriPhy is an auditable physical‑verification system that evaluates generated video by compiling prompts into typed physical obligations and a statically validated execution plan before any frames are observed. During execution, it gates calls to frozen low‑level experts (e.g., segmentation, tracking, counting, depth, OCR, audio‑event detection) and returns provenance‑carrying evidence records, which are mapped to a three‑valued state (supported, contradicted, unknown) with full traceability. On a 1,500‑clip corpus of human‑annotated flaw records, VeriPhy accounts for 228 failures out of 304, outperforming a published question‑decomposition evaluator that accounts for 164, while also providing auditable evidence for each verdict.
The paper investigates when a language‑model judge can truly ground its verdicts in code correctness. It shows that current multi‑agent verification methods rely on evidence that is both independent of the answer and distinct between candidates—conditions that fail in code judging. By analyzing two label‑free measurements from the judge’s logs, the authors demonstrate that gating on one measurement allows the system to decline uncertain comparisons, improving accuracy from 20.7% to 36.9% while still answering half of all cases.
By Salma Roshdy Aly, Hussein Assaf, Ziad Kobti
arXiv:2509. 00761v4 Announce Type: replace Abstract: Large language models are increasingly deployed for legal question answering, where evaluations typically focus on multiple-choice accuracy.
By Boqin Yuan, Ziqi Wang
arXiv:2608. 12877v1 Announce Type: new Abstract: Multi-hop fact verification, which verifies claims by reasoning over multiple pieces of evidence, is critical for combating misinformation on social media yet remains highly challenging.
By Runze Zhao, Zixin Tang, Xiaoshuai Hao, Leyuan Chang, Xiaopeng Fu, Boyu Qiao, Dongyang Zhang
VeriPhy is an auditable physical‑verification system that transforms a text prompt into typed physical obligations and a statically validated execution plan before any video frames are generated. During execution, it gates calls to frozen low‑level experts (segmentation, tracking, counting, depth, OCR, audio‑event detection, etc.) and records provenance‑carrying evidence for each action. The system maps these records to a three‑valued state—supported, contradicted, or unknown—providing traceable verdicts that can be used to refine generation models.
By Wenzhuo Xu, Yuchen Zhu, Chongjian Ge, Xuan Shen, Jing Shi, Jason Kuen, Yongxin Chen, Molei Tao, Christopher McComb, Noelia Grande Guti\'errez, Jiuxiang Gu
arXiv:2608.25937v2 Announce Type: replace
Abstract: Multi-agent systems (MAS) sometimes already have the potential to answer correctly, but still report a wrong answer. Explaining this outcome is dif...
By Jia-Hao Ji, Sijie Li, Jiabei Cheng, Zixi She, Jin-Tai Yu, Zhiyuan Yuan
arXiv:2609.10293v1 Announce Type: new
Abstract: In high-stakes domains such as legal practice, a language-model answer is only useful to the extent that a reader can verify each claim against the sou...
By Chen Qian, Yimeng Wang, Yu Chen, Lingfei Wu, Andreas Stathopoulos
LabourCrew is a multi‑agent Retrieval‑Augmented Generation (RAG) framework designed for trustworthy statutory question answering in labour law. It introduces three grounding mechanisms: StatuteGraph, an evidence‑exchange ledger, and a calibrated trust gate that controls false‑accept rates. Evaluated on a Bangla Labour Act QA set, LabourCrew achieves a false‑accept rate of 0.081 and higher answer relevancy than existing RAG methods, demonstrating that calibrated abstention is key to auditable legal QA.
By Fatema Tuj Johora Faria, Mukaffi Bin Moin, Jubayer Al Mahmud, M. F. Mridha, Md. Alam Hossain
The paper "trajectory-judge: What Outcome-Only LLM Judges Miss on Agent Trajectories" examines the limitations of outcome-only evaluation for large language model agents. Using a deterministic tool‑using support‑desk environment with a scripted oracle policy and a fault injector, the authors compare five different judging approaches—programmatic rules, outcome‑only, step‑rubric at two model sizes, and a self‑consistency ensemble—on metrics such as detection, step localisation, fault typing, calibration, and cost across 400 trajectories. The study finds that outcome‑only judges miss many silent faults and generate false positives, while step‑rubric judges achieve higher recall with no false alarms but at greater cost, and that none of the judges read the final reply, allowing fabricated promises to evade detection.
"whyItMatters":"The findings highlight that current production‑default outcome‑only evaluations can overlook critical failures in agent behavior, underscoring the need for more nuanced, step‑level judging methods to ensure reliable LLM agent performance."
By Hadi Mohammadi