arXiv:2606. 16603v1 Announce Type: cross Abstract: LLM-based agents have demonstrated strong capabilities in data-intensive analytical tasks, yet their outputs are rarely verifiable: a reliance on linear text trajectories makes their reasoning difficult to audit.
By Jiajie Jin, Zhao Yang, Wenle Liao, Yuyang Hu, Guanting Dong, Xiaoxi Li, Yutao Zhu, Zhicheng Dou
HyGRAIL is a framework for discovering scientific hypotheses in incomplete knowledge graphs by combining a graph neural network (GNN) triage with large language model (LLM) review. The GNN scores candidate hypotheses and routes only ambiguous cases to the LLM, which receives structured evidence from the graph converted into natural language. Experiments on MatKG show HyGRAIL achieves the highest F1 score, improves over baselines, and cuts LLM calls by over 54%.
By Yihang Sun, Zhihan Zhu, Zhiyuan Jiang, Jingyi Ge, Zixuan Li, Jiaxuan You
Scientific knowledge graphs organize entities and relations extracted from scientific literature, but they remain inherently incomplete. Missing typed links in such graphs can therefore represent plau...
GraphCert introduces a method to bootstrap graph reasoning agents by generating graph‑grounded question‑answer pairs and certifying the supporting evidence. The approach uses a Bootstrapped Graph Quizzer to produce QA pairs, then executes and semantically curates the evidence into certified rubrics that guide reward‑based training of a Graph Solver. Experiments on five GRBENCH domains show GraphCert outperforms larger LLM agents and demonstrates robust policy transfer across heterogeneous graphs.
By Weiqi Jiang, Yuchen Ying, Rui Wang, Kaixuan Chen, Bingde Hu, Shunyu Liu, Yu Wang, Tongya Zheng
arXiv:2607. 07492v1 Announce Type: new Abstract: Many reasoning tasks are not well described by a single left-to-right chain: a solver may need to pursue a plausible branch, observe delayed failure, and return to the latest prefix that can still be completed.
By Dmitry Beresnev, Vladimir Makharev, Roman Khalikov, Ivan Oseledets, Petr Anokhin
The paper introduces Stochastic Semantic Evidence Graphs (SSEGs), a hierarchical stochastic directed acyclic graph that models uncertainty in AI-agent workflows, from evidence and retrieval to generation and decision mapping. SSEGs expand language nodes into autoregressive token subgraphs, optionally apply semantic reduction and calibration, and preserve uncertain claim–passage relations while propagating Fréchet bounds. The authors derive pathwise error bounds, use nodewise terms to trigger governance checks, and demonstrate through experiments that SSEGs can detect and quantify where uncertainty enters and propagates in AI outputs.
By Matthew Francis Dixon
arXiv:2609.27297v2 Announce Type: replace
Abstract: Agentic science envisions many autonomous agents investigating concurrently while building on a shared, evolving body of scientific knowledge. This...
By Yuan Huang, Sihan Hu, Hongyu Gu, Chao Ma, Jiaxing Zhang, Zhiyong Zou, Caiyu Fan, Yan Xiao, Mingjun Xu, Chenyu Xie, Mingzhen Ju, Zhehao Ma, Qi Zhang, Baozong Wang, Yu Li, Zhiyuan Yao, Ruoxue Liao, Xinyu Li, Linfeng Zhang, Kun Chen, Weinan E
The paper introduces a typed reasoning graph framework to compare human and large language model (LLM) reasoning paths in scientific fact‑checking. By modeling explanations as graphs linking false claims to study context, findings, premises, and fallacy labels, the authors enable one‑to‑one alignment of human and LLM reasoning at the sub‑graph level. Using 84 false claims from MISSCIPLUS, they evaluate GPT‑5, Claude Opus 4.7, and Qwen3‑32B, finding distinct performance patterns: Qwen3‑32B has the lowest verdict failure rate, GPT‑5 shows the highest human alignment, and Claude Opus 4.7, while weak at verdict prediction, often produces valid reasoning in successful cases.
By Abdul Ghafoor, Muhammad Arslan Manzoor, Yufang Hou
arXiv:2608. 04738v1 Announce Type: new Abstract: Autonomous research agents can generate hypotheses, execute experiments, and draft manuscripts, yet their outputs often contain unsupported claims and inconsistencies between research questions, experiments, results, and conclusions.
By Zhenjiang Ren, Ruiji Li, Xujing Zhang, Ziliang Pang, Shuo Ren, Jiajun Zhang
LEDGER is a tracing and review system for large language model agents that constructs layered trace graphs from observed sessions. It groups raw trace records into Evidence Nodes and Workflow Nodes, anchors artifacts as evidence, and adds typed semantic edges linking claims to supporting actions, artifacts, and checks. The resulting traces reveal workflow decisions, artifact lineage, repair steps, validation coverage, and claim‑support paths for evidence‑centered audit.
By Daehong Kim, Haichao Miao, Shusen Liu
GraphEcho is a benchmark that examines how large language model agents navigate graph paths and handle evidence redundancy. It tests whether agents treat repeated encounters as additional corroboration by varying path counts and evidential origins while keeping evidence content constant. The study finds that redundant paths increase repeated walks, and that provenance-aware post‑training can reduce revisits but may limit source diversity, revealing a gap between efficient exploration and effective evidence use.
By Sikun Wang, Yixi Zhou, Lei Fan, Fan Zhang
arXiv:2608.29612v1 Announce Type: new
Abstract: Sustained scientific work requires a knowledge substrate that carries interpretation across tasks and preserves paths to source evidence. We call this...
By Shi-Ju Ran, Kun Zhang, Xi Wu, Liu-Si Yang, Wen-Jun Li