arXiv Computation and Language

PRISM-BN: A Controlled Corpus and Benchmark for Text-to-Parameterized Bayesian Network Extraction

arXiv Machine Learning
Sep 15

Scalable partial information decomposition for symptom networks via supervised embeddings

The paper introduces ePID, an embedding-based approach that scales partial information decomposition (PID) to large symptom networks by compressing non‑focal symptoms into a low‑cardinality discrete embedding. Using a supervised Agglomerative Conditional Information Bottleneck (ACIB) embedding, ePID accurately recovers source‑unique, remainder‑unique, redundant, and synergistic components for each ordered source‑target pair across 83 real‑world datasets, outperforming 12 other embeddings. Applied to PHQ‑9 and the Interpersonal Reactivity Index, ePID reveals distinct patterns of redundancy and synergy that align with each instrument’s construction, demonstrating its ability to separate overlapping from interaction‑dependent information in symptom networks.

By Cillian Hourican, Eric Dignum, Rick Quax, Debraj Roy
arXiv AI
Sep 17

Extracting Probabilistic Knowledge from Large Language Models for Bayesian Network Parameterization

The paper investigates how Large Language Models can be used to approximate domain expert priors for Bayesian Networks by extracting probabilistic knowledge about real‑world events. Experiments on eighty publicly available networks across domains such as healthcare and finance show that LLM‑derived conditional probabilities outperform random, uniform, and next‑token baselines. The authors also demonstrate that these LLM‑generated priors can refine data‑driven distributions, especially when data is scarce, and provide the first comprehensive baseline for evaluating LLM performance in probabilistic knowledge extraction.

By Aliakbar Nafar, Kristen Brent Venable, Zijun Cui, Parisa Kordjamshidi
arXiv AI
Jul 28

Structure Over Scale: Schema-Constrained Causal Graphs for RAG

arXiv:2607. 22592v1 Announce Type: new Abstract: Graph-based retrieval-augmented generation (GraphRAG) grounds answers in structured knowledge, but current systems extract entities and relationships exhaustively, producing graphs whose size and construction cost scale with corpus length rather than with the reasoning a query requires.

By Marc Saouda (Boston Consulting Group), Rajprakash Bale (Boston Consulting Group), Eren Aldis (Boston Consulting Group), Cloves Almeida (Boston Consulting Group)
arXiv AI
Sep 21

Auditing a KB Elicitation of Frontier LLM Knowledge: A Multi-dimensional Analysis of GPTKB v1.5

The paper presents a framework and results of a multi-dimensional analysis of GPTKB v1.5, a 100‑million‑fact knowledge base elicited from GPT‑4.1. It shows that the LLM’s factual knowledge differs markedly from established knowledge bases and that its accuracy is lower than suggested by prior benchmarks. The study also identifies inconsistency, ambiguity, and hallucinations as major issues, pointing to future research directions in neuro‑symbolic AI for extracting, consolidating, and verifying factual LLM knowledge.

By Shrestha Ghosh, Luca Giordano, Yujia Hu, Tuan-Phong Nguyen, Simon Razniewski