arXiv AI

GraphEcho: Structural Redundancy and Evidence Provenance in LLM Graph Agents

GraphEcho is a benchmark that examines how large language model agents navigate graph paths and handle evidence redundancy. It tests whether agents treat repeated encounters as additional corroboration by varying path counts and evidential origins while keeping evidence content constant. The study finds that redundant paths increase repeated walks, and that provenance-aware post‑training can reduce revisits but may limit source diversity, revealing a gap between efficient exploration and effective evidence use.

arXiv AI
Aug 6

EviGraph: Evidence-Guided Autonomous Research Agents

arXiv:2608. 04738v1 Announce Type: new Abstract: Autonomous research agents can generate hypotheses, execute experiments, and draft manuscripts, yet their outputs often contain unsupported claims and inconsistencies between research questions, experiments, results, and conclusions.

By Zhenjiang Ren, Ruiji Li, Xujing Zhang, Ziliang Pang, Shuo Ren, Jiajun Zhang
arXiv Computation and Language
Sep 25

Stochastic Semantic Evidence Graphs: Uncertainty Propagation and Governance for Agentic AI

The paper introduces Stochastic Semantic Evidence Graphs (SSEGs), a hierarchical stochastic directed acyclic graph that models uncertainty in AI-agent workflows, from evidence and retrieval to generation and decision mapping. SSEGs expand language nodes into autoregressive token subgraphs, optionally apply semantic reduction and calibration, and preserve uncertain claim–passage relations while propagating Fréchet bounds. The authors derive pathwise error bounds, use nodewise terms to trigger governance checks, and demonstrate through experiments that SSEGs can detect and quantify where uncertainty enters and propagates in AI outputs.

By Matthew Francis Dixon
arXiv AI
3d ago

GraphCert: Bootstrap Agentic Graph Reasoning with Certified Evidence Rubrics

GraphCert introduces a method to bootstrap graph reasoning agents by generating graph‑grounded question‑answer pairs and certifying the supporting evidence. The approach uses a Bootstrapped Graph Quizzer to produce QA pairs, then executes and semantically curates the evidence into certified rubrics that guide reward‑based training of a Graph Solver. Experiments on five GRBENCH domains show GraphCert outperforms larger LLM agents and demonstrates robust policy transfer across heterogeneous graphs.

By Weiqi Jiang, Yuchen Ying, Rui Wang, Kaixuan Chen, Bingde Hu, Shunyu Liu, Yu Wang, Tongya Zheng
arXiv AI
3d ago

PathAnchor: Path-Structured Evidence for Scientific Agents

PathAnchor is a new scientific reasoning system that uses path-structured evidence workspaces instead of independent passages or concepts. It retrieves source-linked Material‑Sensor‑Signal‑System trajectories that preserve role, direction, and supporting evidence, and a controller uses read‑only tools to search, trace, and open exact evidence before producing a claim‑cited answer. In evaluations on 120 flexible‑sensor questions, PathAnchor achieved an 82.6% score, outperformed six other systems, and improved source recall, citation completeness, and reduced tool calls compared to unordered concept graphs.

By Qiuhui Chen, Jiafan Lu, Shuaimin Tang, Tao Dai, Suyuan Wang, Chenrui Ji, Zhenglei Zhou, Weimin Zhong
arXiv Computation and Language
Sep 3

HyGRAIL: Cost-Aware and Evidence-Grounded Scientific Hypothesis Discovery over Knowledge Graphs

HyGRAIL is a framework for discovering scientific hypotheses in incomplete knowledge graphs by combining a graph neural network (GNN) triage with large language model (LLM) review. The GNN scores candidate hypotheses and routes only ambiguous cases to the LLM, which receives structured evidence from the graph converted into natural language. Experiments on MatKG show HyGRAIL achieves the highest F1 score, improves over baselines, and cuts LLM calls by over 54%.

By Yihang Sun, Zhihan Zhu, Zhiyuan Jiang, Jingyi Ge, Zixuan Li, Jiaxuan You
arXiv AI
Sep 7

TruthInsightBench: An Evidence-Grounded Benchmark for Automated Evaluation of Open-Ended Scientific Discovery Agents

TruthInsightBench is a new benchmark designed to evaluate automated scientific discovery agents by presenting them with 40 blind tasks drawn from peer‑reviewed studies across ten domains. Each task provides only a neutral objective and frozen data, withholding source conclusions, expected values, and analysis paths, forcing agents to determine which claim the data support. A fixed LLM‑based judge scores agents on evidentiary maturity across six dimensions, using 29 artifact‑grounded items, enabling fully automated, repeatable evaluation without human grading.

By Zhibo Yang, Chen Zhang, Yuewei Zhang, Hao Wang
arXiv AI
Jul 14

FIRE-Bench: Evaluating AI Agents on the Rediscovery of Scientific Insights

arXiv:2602. 02905v2 Announce Type: replace Abstract: Autonomous agents powered by large language models (LLMs) promise to accelerate scientific discovery end-to-end, but rigorously evaluating their capacity for verifiable discovery remains a central challenge.

By Zhen Wang, Fan Bai, Zhongyan Luo, Jinyan Su, Kaiser Sun, Xinle Yu, Jieyuan Liu, Kun Zhou, Claire Cardie, Mark Dredze, Zhiting Hu, Eric P. Xing
arXiv AI
Sep 12

MOSAIC: Query-Aware Exploration Policy Adaptation for GraphRAG

MOSAIC is a training‑free framework that adapts Graph Retrieval‑Augmented Generation (GraphRAG) to each query by converting query‑specific evidence needs into a bounded policy over seed selection, traversal, stopping, and evidence selection. It keeps the corpus graph, indexes, scoring, grounding, and answer generation shared, while an LLM analyzer tailors the exploration strategy per query. On GraphRAG‑Bench, MOSAIC improves answer correctness by over 5 points on Medical and 4 points on Novel, achieves high evidence recall and context relevancy, and reduces path and evidence evaluations compared to fixed policies.

By EunKyeong Lee, Kyeong-Jin Oh, Jinwon Kim, Hye Woo Lee, Minsang Song, Hyeongjun Jang, Junyoung Youn