arXiv Machine Learning

Exactness at Inference: A Representational Criterion for Out-of-Distribution Generalization

arXiv AI
Aug 19

Neuro-symbolic learning over OWL 2 DL via consequence-based compilation to differentiable circuits

Baobab compiles an OWL 2 DL (ΣROIQ) ontology with a finite ABox into a Sentential Decision Diagram (SDD), saturating a propositional core and instantiating remaining DL features over the active domain. The resulting evidence‑conditioned weighted model count trains a perception network to recognize real images under partial ABox supervision, enabling a CNN to recover latent ontology concepts that an independent perception would miss. When supervision allows multiple ontology‑consistent completions, Baobab’s mixture indexed by query justifications represents the calibrated posterior, achieving Bayes‑optimal performance on a real‑image MNIST task where single‑WMC and learned mixtures fail, thereby characterizing and mitigating reasoning shortcuts in a non‑Horn description logic.

By Olga Mashkova, Asaad Mohammedsaleh, Fernando Zhapa-Camacho, Robert Hoehndorf
arXiv AI
Sep 15

Certifiably Interpretable Training of ReLU-MLPs for Boolean Tasks with Guaranteed Truth-Table Generalization

The paper introduces MACCHIATO, a training algorithm that builds a ReLU‑MLP from partial truth‑table data while simultaneously constructing an explicit Boolean circuit over AND, OR, and XOR gates that certifies the network’s computation. The method iteratively projects residuals onto low‑dimensional Boolean classes, compiles the resulting circuit into a ReLU‑MLP, and uses logic minimization and influence‑based variable selection to achieve a six‑layer network with provable truth‑table error bounds. Experiments on synthetic random‑junta tasks show that these certified networks outperform Adam‑trained MLPs in data‑sparse or projection‑aligned regimes and complete faster than flat ESPRESSO in certain settings.

By Hrad Ghoukasian, Anastasis Kratsios
arXiv AI
Sep 7

DODR: Deterministic Operator-Driven Reasoning in Latent Space

The paper introduces DODR, a deterministic operator‑driven reasoning architecture that models reasoning as graph computations in a high‑dimensional linear‑algebraic space, replacing token‑level sampling with matrix operations. Reasoning states are snapshot vectors of semantic units, and three trainable matrix operators—deduction, induction, and abduction—implement Peirce’s inference types. Experiments on 503 records demonstrate near‑perfect deduction, high generalization for induction, and significant gains for abduction, while the design guarantees zero hallucination and supports continual learning.

By Weicai Huang (Beijing MQPat Technologies, Co., Ltd.)
arXiv AI
Sep 21

Large Language Models As Shannon Lossy Compressors Not Solomonoff Induction Estimators: The Singularity Is Not Near Without Symbolic Model Synthesis

The paper argues that Large Language Models (LLMs) do not function as Solomonoff induction estimators because their training objectives—cross‑entropy, negative log‑likelihood, and next‑token prediction—optimize fit to a supplied conditional distribution rather than a program‑weighted universal mixture. It further contends that additional computation alone does not transform these models into optimal predictors without external hyper‑parameter or architectural changes. The authors suggest that neurosymbolic machine learning, exemplified by models such as Fable and Astra, represents a shift toward symbolic model synthesis, moving beyond purely statistical LLMs.

By Hector Zenil, Abicumaran Uthamacumaran, Luan Ozelim