arXiv Machine Learning

Uncertainty in Representation Learning on Knowledge Graphs

arXiv AI
Jul 10

Can We Trust LLM's Logic? Quantifying Uncertainty, Coherence, and Robustness via a Graph-Based Framework

arXiv:2607. 08017v1 Announce Type: cross Abstract: Large-Language Models (LLMs) can be prone to flawed and unfaithful reasoning that decoding strategies like Self-Consistency (SC) fail to detect as they evaluate only final-answer agreement while ignoring the logical validity of intermediate steps.

By Riccardo Revalor, Jalees Rehman, Debjit Pal
arXiv AI
Sep 3

Spectral Initialization and Scheduled Graph Smoothness for Uncertain Knowledge Graph Completion

The paper introduces QUEST, a method for uncertain knowledge graph completion that adds no trainable parameters to the standard pipeline. QUEST first initializes entity embeddings using the smallest non‑trivial eigenvectors of the confidence‑weighted graph Laplacian, thereby preserving community and hub structure before training. It then applies an unbiased mini‑batch Dirichlet energy regularizer to enforce early‑stage structural consistency, leading to improved confidence and link prediction on most metric‑dataset pairs and eliminating instability spikes on dense graphs.

By Md Abrar Jahin, Taufikur Rahman Fuad, Jay Pujara, Craig A. Knoblock
arXiv AI
Sep 7

GUT: Quantifying and Optimizing the Reasoning Uncertainty of LLMs via Graph Complexity

The paper introduces GUT, a method that uses directed acyclic graphs to represent all possible reasoning branches of Large Language Models (LLMs). It comprises two modules: GUT-Q, which quantifies reasoning uncertainty by approximating graph complexity, and GUT-O, which reduces uncertainty through reinforcement learning that rewards lower uncertainty. Experiments on four LLMs across five datasets demonstrate GUT’s effectiveness in measuring and mitigating reasoning uncertainty.

By Shuang Liang, Xin-Yu Hu, Xiang-Jun Ou, Shao-Qun Zhang
arXiv Machine Learning
Jun 19

Quantifying Aleatoric Uncertainty of In-Context Learning for Robust Measure of LLM Prediction Confidence

arXiv:2606. 19353v1 Announce Type: cross Abstract: In-Context Learning (ICL) allows LLMs to adapt to new tasks from a few demonstrations, but its reliability remains a concern: predictions are highly sensitive to both prompt design and the model's ability to understand the context, obscuring whether failures arise from data properties or model limitations.

By Jinseok Chung, Minkyoung Song, Hyunji Jung, Namhoon Lee
arXiv AI
Sep 24

A hierarchy of faithfulness criteria for knowledge base completion

The paper introduces a hierarchy of four increasingly strict faithfulness criteria—discrimination, logical admissibility, monotonic logical faithfulness, and probabilistic logical faithfulness—for evaluating knowledge base completion models, particularly when the target is a description logic knowledge base. It demonstrates that ranking accuracy alone does not guarantee logical faithfulness and shows that current embedding models fail to satisfy any of the criteria across the hierarchy. The authors provide a formal grounding for the strongest criterion using relative model counts and evaluate several models on εL ontologies, revealing gaps between performance metrics and logical correctness.

By Olga Mashkova, Robert Hoehndorf