arXiv AI

Provenance-Enhanced Statements in Knowledge Graphs

arXiv:2606. 15246v1 Announce Type: cross Abstract: Provenance-enhanced statements of the form "according to $X$, $\varphi$" are pervasive in contemporary knowledge graphs, especially in domains where graph content primarily represents claims, interpretations, and hypotheses (\emph{capta}) rather than observer-independent facts (\emph{data}).

arXiv AI
Jun 16

VeriGraph: Towards Verifiable Data-Analytic Agents

arXiv:2606. 16603v1 Announce Type: cross Abstract: LLM-based agents have demonstrated strong capabilities in data-intensive analytical tasks, yet their outputs are rarely verifiable: a reliance on linear text trajectories makes their reasoning difficult to audit.

By Jiajie Jin, Zhao Yang, Wenle Liao, Yuyang Hu, Guanting Dong, Xiaoxi Li, Yutao Zhu, Zhicheng Dou
arXiv AI
Aug 25

Walking on the DARKSIDE

arXiv:2608.23370v1 Announce Type: new Abstract: Large Language Models (LLMs) recognise patterns but do not natively track the path of exclusions that a coherent discourse demands. When an input rests...

By Aldo Gangemi, Emanuele Bottazzi
arXiv AI
Sep 4

Semantic Bayesian World Models

Semantic Bayesian World Models (SBWMs) propose a shift from static knowledge graphs to a dynamic, probabilistic fabric of beliefs that can be updated via Bayesian conditioning and influenced by actions. The approach aims to bridge the gap between crisp factual assertions and the probabilistic reasoning of foundation models and autonomous agents, enabling richer inference in scenarios such as home‑security decisions, actuarial estimates, and planning tasks. Realizing SBWMs requires new tools for belief annotation, probabilistic entailment, semantic calibration, and protocols for belief exchange among agents.

By Tommaso Soru
arXiv AI
Sep 7

GRACE: Graph-Grounded Reflective Agent Copilot Engine for Expert-in-the-Loop Knowledge Expansion

The paper introduces GRACE, a framework that breaks down large language model (LLM) responses into atomic claims and grounds them against trusted knowledge priors using a weighted bipartite graph. Edge weights enable weighted centrality analysis to classify claims as Grounded, Refuted, or Boundary, identifying hallucinations and frontier knowledge. An objective called Return on Attention (RoA) prioritizes expert review only for high‑uncertainty claims, and verified claims become new evidence anchors, creating a loop that expands the knowledge base across iterations.

By John Seon Keun Yi, Joshua R. Minot, Dokyun Lee
arXiv AI
Aug 25

Beyond Verdicts: A Graph-Based Analysis of Human and LLM Reasoning in Scientific Fact-Checking

The paper introduces a typed reasoning graph framework to compare human and large language model (LLM) reasoning paths in scientific fact‑checking. By modeling explanations as graphs linking false claims to study context, findings, premises, and fallacy labels, the authors enable one‑to‑one alignment of human and LLM reasoning at the sub‑graph level. Using 84 false claims from MISSCIPLUS, they evaluate GPT‑5, Claude Opus 4.7, and Qwen3‑32B, finding distinct performance patterns: Qwen3‑32B has the lowest verdict failure rate, GPT‑5 shows the highest human alignment, and Claude Opus 4.7, while weak at verdict prediction, often produces valid reasoning in successful cases.

By Abdul Ghafoor, Muhammad Arslan Manzoor, Yufang Hou