arXiv Computation and Language By Matthew Francis Dixon

Stochastic Semantic Evidence Graphs: Uncertainty Propagation and Governance for Agentic AI

Read the original on arXiv Computation and Language →

The paper introduces Stochastic Semantic Evidence Graphs (SSEGs), a hierarchical stochastic directed acyclic graph that models uncertainty in AI-agent workflows, from evidence and retrieval to generation and decision mapping. SSEGs expand language nodes into autoregressive token subgraphs, optionally apply semantic reduction and calibration, and preserve uncertain claim–passage relations while propagating Fréchet bounds. The authors derive pathwise error bounds, use nodewise terms to trigger governance checks, and demonstrate through experiments that SSEGs can detect and quantify where uncertainty enters and propagates in AI outputs.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Jun 16

VeriGraph: Towards Verifiable Data-Analytic Agents

arXiv:2606. 16603v1 Announce Type: cross Abstract: LLM-based agents have demonstrated strong capabilities in data-intensive analytical tasks, yet their outputs are rarely verifiable: a reliance on linear text trajectories makes their reasoning difficult to audit.

By Jiajie Jin, Zhao Yang, Wenle Liao, Yuyang Hu, Guanting Dong, Xiaoxi Li, Yutao Zhu, Zhicheng Dou
arXiv AI
Sep 17

GraphEcho: Structural Redundancy and Evidence Provenance in LLM Graph Agents

GraphEcho is a benchmark that examines how large language model agents navigate graph paths and handle evidence redundancy. It tests whether agents treat repeated encounters as additional corroboration by varying path counts and evidential origins while keeping evidence content constant. The study finds that redundant paths increase repeated walks, and that provenance-aware post‑training can reduce revisits but may limit source diversity, revealing a gap between efficient exploration and effective evidence use.

By Sikun Wang, Yixi Zhou, Lei Fan, Fan Zhang
arXiv AI
Aug 6

EviGraph: Evidence-Guided Autonomous Research Agents

arXiv:2608. 04738v1 Announce Type: new Abstract: Autonomous research agents can generate hypotheses, execute experiments, and draft manuscripts, yet their outputs often contain unsupported claims and inconsistencies between research questions, experiments, results, and conclusions.

By Zhenjiang Ren, Ruiji Li, Xujing Zhang, Ziliang Pang, Shuo Ren, Jiajun Zhang
arXiv AI
Sep 7

GRACE: Graph-Grounded Reflective Agent Copilot Engine for Expert-in-the-Loop Knowledge Expansion

The paper introduces GRACE, a framework that breaks down large language model (LLM) responses into atomic claims and grounds them against trusted knowledge priors using a weighted bipartite graph. Edge weights enable weighted centrality analysis to classify claims as Grounded, Refuted, or Boundary, identifying hallucinations and frontier knowledge. An objective called Return on Attention (RoA) prioritizes expert review only for high‑uncertainty claims, and verified claims become new evidence anchors, creating a loop that expands the knowledge base across iterations.

By John Seon Keun Yi, Joshua R. Minot, Dokyun Lee
arXiv AI
Aug 20

LEDGER: Claim-to-Evidence Trace Graphs for Auditing LLM Agents

LEDGER is a tracing and review system for large language model agents that constructs layered trace graphs from observed sessions. It groups raw trace records into Evidence Nodes and Workflow Nodes, anchors artifacts as evidence, and adds typed semantic edges linking claims to supporting actions, artifacts, and checks. The resulting traces reveal workflow decisions, artifact lineage, repair steps, validation coverage, and claim‑support paths for evidence‑centered audit.

By Daehong Kim, Haichao Miao, Shusen Liu