The paper introduces ePID, an embedding-based approach that scales partial information decomposition (PID) to large symptom networks by compressing non‑focal symptoms into a low‑cardinality discrete embedding. Using a supervised Agglomerative Conditional Information Bottleneck (ACIB) embedding, ePID accurately recovers source‑unique, remainder‑unique, redundant, and synergistic components for each ordered source‑target pair across 83 real‑world datasets, outperforming 12 other embeddings. Applied to PHQ‑9 and the Interpersonal Reactivity Index, ePID reveals distinct patterns of redundancy and synergy that align with each instrument’s construction, demonstrating its ability to separate overlapping from interaction‑dependent information in symptom networks.
By Cillian Hourican, Eric Dignum, Rick Quax, Debraj Roy
The paper investigates how Large Language Models can be used to approximate domain expert priors for Bayesian Networks by extracting probabilistic knowledge about real‑world events. Experiments on eighty publicly available networks across domains such as healthcare and finance show that LLM‑derived conditional probabilities outperform random, uniform, and next‑token baselines. The authors also demonstrate that these LLM‑generated priors can refine data‑driven distributions, especially when data is scarce, and provide the first comprehensive baseline for evaluating LLM performance in probabilistic knowledge extraction.
By Aliakbar Nafar, Kristen Brent Venable, Zijun Cui, Parisa Kordjamshidi
arXiv:2609.15007v1 Announce Type: cross
Abstract: Large language models are increasingly used as natural-language interfaces to structured data, yet they remain unreliable when answers require consis...
By Jackson Hassell, Chen Shen, Estevam Hruschka
arXiv:2608. 03868v1 Announce Type: cross Abstract: Causal Discovery (CD) from observational data faces two fundamental challenges.
By Abhinav Thorat, Ravi Kumar Kolla, Vishak K Bhat, Harsh Vardhan Singh Chauhan, Niranjan Pedanekar
arXiv:2607. 14314v1 Announce Type: new Abstract: Seizure diagnosis from EEG signals is a critical yet persistently challenging task, due to the complicated neural dynamics and the spurious connections in inter-channel modeling.
By Lincan Li, Zheng Chen, Yushun Dong
arXiv:2601.01368v2 Announce Type: replace
Abstract: Score-based causal discovery in the presence of unobserved confounders requires both a consistent scoring criterion and an efficient search over gr...
By Mujin Zhou, Ignavier Ng, Junzhe Zhang
arXiv:2608. 12640v1 Announce Type: cross Abstract: Causal discovery aims to uncover the underlying causal relationships given data generated from a system.
By Cixuan Zhang, Guy Van den Broeck, Benjie Wang
arXiv:2608. 04930v1 Announce Type: cross Abstract: Bayesian causal discovery seeks to determine the posterior distribution of causal theories, which are interpreted as directed acyclic graphs (DAGs) that explain the observed data.
By Shrenik Zinage
arXiv:2606. 24488v1 Announce Type: cross Abstract: Learning causal models from fragmented biomedical data is challenging because clinical, molecular, and imaging variables are often incomplete or not jointly observed.
By Inam Ullah, Imran Razzak, Shoaib Jameel
arXiv:2606. 11831v1 Announce Type: cross Abstract: Neural relational inference (NRI) methods discover interaction graphs from trajectories through variational reasoning on discrete potential edges.
By Qi Shao, Hao Guo, Jiawen Chen, Duxin Chen, Wenwu Yu
arXiv:2607. 22592v1 Announce Type: new Abstract: Graph-based retrieval-augmented generation (GraphRAG) grounds answers in structured knowledge, but current systems extract entities and relationships exhaustively, producing graphs whose size and construction cost scale with corpus length rather than with the reasoning a query requires.
By Marc Saouda (Boston Consulting Group), Rajprakash Bale (Boston Consulting Group), Eren Aldis (Boston Consulting Group), Cloves Almeida (Boston Consulting Group)
The paper presents a framework and results of a multi-dimensional analysis of GPTKB v1.5, a 100‑million‑fact knowledge base elicited from GPT‑4.1. It shows that the LLM’s factual knowledge differs markedly from established knowledge bases and that its accuracy is lower than suggested by prior benchmarks. The study also identifies inconsistency, ambiguity, and hallucinations as major issues, pointing to future research directions in neuro‑symbolic AI for extracting, consolidating, and verifying factual LLM knowledge.
By Shrestha Ghosh, Luca Giordano, Yujia Hu, Tuan-Phong Nguyen, Simon Razniewski