arXiv AI By Olga Mashkova, Robert Hoehndorf

A hierarchy of faithfulness criteria for knowledge base completion

Read the original on arXiv AI →

The paper introduces a hierarchy of four increasingly strict faithfulness criteria—discrimination, logical admissibility, monotonic logical faithfulness, and probabilistic logical faithfulness—for evaluating knowledge base completion models, particularly when the target is a description logic knowledge base. It demonstrates that ranking accuracy alone does not guarantee logical faithfulness and shows that current embedding models fail to satisfy any of the criteria across the hierarchy. The authors provide a formal grounding for the strongest criterion using relative model counts and evaluate several models on εL ontologies, revealing gaps between performance metrics and logical correctness.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jul 10

Can We Trust LLM's Logic? Quantifying Uncertainty, Coherence, and Robustness via a Graph-Based Framework

arXiv:2607. 08017v1 Announce Type: cross Abstract: Large-Language Models (LLMs) can be prone to flawed and unfaithful reasoning that decoding strategies like Self-Consistency (SC) fail to detect as they evaluate only final-answer agreement while ignoring the logical validity of intermediate steps.

By Riccardo Revalor, Jalees Rehman, Debjit Pal
arXiv AI
2d ago

Ontology-Grounded, Reasoner-Verified Benchmarks for Evaluating LLM Reasoning in Scientific AI

The paper introduces a pipeline that automatically creates ontology‑grounded multiple‑choice question benchmarks for evaluating large language models (LLMs) on logical reasoning tasks in scientific AI. By using OWL 2 ontologies, correct answers are guaranteed by design and distractors are generated and formally verified as incorrect through an OWL reasoner. Experiments on three ontologies—Pizza, PMDco, and DOID—yielded 112, 2,491, and 15,216 MCQs, respectively, with high natural‑language quality and challenging zero‑shot performance for six LLMs.

By Nishtha N. Vaidya, Stephan Grimm, Thomas Hubauer, Thomas A. Runkler
arXiv AI
Jun 2

KnowledgeBerg: Evaluating Systematic Knowledge Coverage and Compositional Reasoning in Large Language Models

arXiv:2604. 17621v2 Announce Type: replace Abstract: Many real-world questions appear deceptively simple yet implicitly demand two capabilities: (i) systematic coverage of a bounded knowledge universe and (ii) compositional set-based reasoning over that universe, a phenomenon we term "the tip of the iceberg.

By Xiao Zhang, Qianru Meng, Yongjian Chen, Yumeng Wang, Johan Bos
arXiv AI
Aug 19

Neuro-symbolic learning over OWL 2 DL via consequence-based compilation to differentiable circuits

Baobab compiles an OWL 2 DL (ΣROIQ) ontology with a finite ABox into a Sentential Decision Diagram (SDD), saturating a propositional core and instantiating remaining DL features over the active domain. The resulting evidence‑conditioned weighted model count trains a perception network to recognize real images under partial ABox supervision, enabling a CNN to recover latent ontology concepts that an independent perception would miss. When supervision allows multiple ontology‑consistent completions, Baobab’s mixture indexed by query justifications represents the calibrated posterior, achieving Bayes‑optimal performance on a real‑image MNIST task where single‑WMC and learned mixtures fail, thereby characterizing and mitigating reasoning shortcuts in a non‑Horn description logic.

By Olga Mashkova, Asaad Mohammedsaleh, Fernando Zhapa-Camacho, Robert Hoehndorf