GroundEval: A Deterministic Replacement for LLM-as-Judge in Stateful Agent Evaluation
arXiv:2606. 22737v2 Announce Type: replace Abstract: Before letting an agent operate over real context, can you prove it used the right evidence?
arXiv:2608. 02011v2 Announce Type: replace Abstract: Agentic retrieval-augmented generation (RAG) systems can fail before evidence-conditioned reasoning is tested: an agent may retrieve candidate snippets but finalize without inspecting them.
arXiv:2606. 22737v2 Announce Type: replace Abstract: Before letting an agent operate over real context, can you prove it used the right evidence?
arXiv:2608. 12877v1 Announce Type: new Abstract: Multi-hop fact verification, which verifies claims by reasoning over multiple pieces of evidence, is critical for combating misinformation on social media yet remains highly challenging.
The paper introduces AgenticRAG-FP, an interventional benchmark designed to attribute causal failures in agentic retrieval‑augmented generation (RAG) systems. By injecting a certified fault at a specified hop and re‑executing the downstream trajectory, the benchmark evaluates whether post‑hoc diagnostics can correctly identify the fault’s location. Experiments on MuSiQue questions show that coverage‑based diagnosis performs well at hop 1 but poorly at later hops, while counterfactual probes reveal varying diagnostic success depending on propagation depth.
REALHOP introduces a behavioral auditing framework to assess multi‑hop reasoning by measuring the Behavioral Necessity Rate (BNR), which quantifies how often removing targeted evidence prevents correct answers. Across five benchmarks, the framework reveals a wide gap between annotated reasoning chains and actual evidence dependence, with panel‑mean BNR ranging from 16.6% to 48.9%. By re‑binding entities, factorizing relations, adding competing paths, and placing evidence at traceable locations, REALHOP raises BNR dramatically—from 27.4% to 94.4% on MuSiQue questions—while maintaining high overall accuracy and improving performance on long‑context tasks.
arXiv:2607. 12267v1 Announce Type: cross Abstract: Language agents that interleave reasoning and tool use degrade sharply as reasoning chains lengthen, even when each individual step is easy.
arXiv:2608. 06128v1 Announce Type: new Abstract: Search agents extend large language models beyond static parametric memory by enabling them to acquire and use ex ternal evidence during multi-step reasoning.
VeriPhy is an auditable physical‑verification system that evaluates generated video by compiling prompts into typed physical obligations and a statically validated execution plan before any frames are observed. During execution, it gates calls to frozen low‑level experts (e.g., segmentation, tracking, counting, depth, OCR, audio‑event detection) and returns provenance‑carrying evidence records, which are mapped to a three‑valued state (supported, contradicted, unknown) with full traceability. On a 1,500‑clip corpus of human‑annotated flaw records, VeriPhy accounts for 228 failures out of 304, outperforming a published question‑decomposition evaluator that accounts for 164, while also providing auditable evidence for each verdict.
arXiv:2509. 23071v2 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) agent development is hindered by the lack of executable ground-truth agent-environment interaction trajectories.
The paper introduces MIRAGE, a benchmark of 750 multi‑step decision tasks designed to test autonomous web agents’ investigative abilities across Wikipedia Forensics, Shopping Admin adjudication, and Reddit Moderation. Each task contains a misleading visible context and a hidden context that holds decisive evidence, allowing performance to be broken down into Investigation, Reasoning, and Decision Accuracy, with an added Investigative Hallucination Rate. Evaluation of eight LLM agents reveals three consistent patterns: agents often reach relevant pages but fail to extract decisive evidence, procedural hints improve investigation but not decision accuracy on Wikipedia tasks, and 12.6% of trajectories include fabricated facts.
arXiv:2606. 00765v1 Announce Type: new Abstract: LLM-based agents increasingly solve complex tasks through long trajectories involving reasoning steps, tool calls, and inter-agent communication.
arXiv:2607. 20952v1 Announce Type: new Abstract: Latent, or silent, reasoning lets language models carry out intermediate computation in continuous vector space instead of words, and is widely assumed to function as an internal scratchpad the model actively consults during inference.
arXiv:2608. 07885v1 Announce Type: new Abstract: Reasoning modes of language models outperform their non-reasoning counterparts on multi-step agentic tasks, but pay a 3-6x premium in output tokens on every episode -- much of it spent re-deriving procedures that are shared across episodes of the same domain.