arXiv Machine Learning

Identifiability and Order-Dimension Limits of In-Context Learning on Partial Orders

arXiv:2608. 14004v1 Announce Type: new Abstract: In-context learning is commonly formalized as inference from examples of a function.

arXiv AI
Jul 24

Representative Sets in Propositional Abduction

arXiv:2607. 21183v1 Announce Type: cross Abstract: The propositional abduction problem is a well-known form of non-monotonic reasoning where we are asked to find an explanation of a given manifestation.

By Johannes Schmidt (J\"onk\"oping University), Mohamed Maizia (J\"onk\"oping University, Link\"oping University), Victor Lagerkvist (Link\"oping University), Johannes K. Fichte (Link\"oping University)
arXiv AI
Jun 2

KnowledgeBerg: Evaluating Systematic Knowledge Coverage and Compositional Reasoning in Large Language Models

arXiv:2604. 17621v2 Announce Type: replace Abstract: Many real-world questions appear deceptively simple yet implicitly demand two capabilities: (i) systematic coverage of a bounded knowledge universe and (ii) compositional set-based reasoning over that universe, a phenomenon we term "the tip of the iceberg.

By Xiao Zhang, Qianru Meng, Yongjian Chen, Yumeng Wang, Johan Bos
arXiv AI
Sep 25

SAGE: Mitigating Long-Horizon Reasoning Biases via Topological Guidance

The paper introduces SAGE, a framework designed to reduce long‑horizon reasoning biases in large language models. It identifies two key biases—exploration bias and compounding bias—arising from complex reasoning spaces and sparse rewards, and proposes Symbolic Closure Analysis (SCA) to understand these effects. SAGE applies algebraic sparsification and hyperbolic structural guidance to suppress spurious branching and provide dense depth‑wise signals, achieving up to an eight‑fold improvement on the Andrews‑Curtis problem across multiple benchmarks and model families.

By Xinyue Zeng, Jiawei Zhang, Yujun Yan, Dawei Zhou
arXiv AI
Sep 7

DODR: Deterministic Operator-Driven Reasoning in Latent Space

The paper introduces DODR, a deterministic operator‑driven reasoning architecture that models reasoning as graph computations in a high‑dimensional linear‑algebraic space, replacing token‑level sampling with matrix operations. Reasoning states are snapshot vectors of semantic units, and three trainable matrix operators—deduction, induction, and abduction—implement Peirce’s inference types. Experiments on 503 records demonstrate near‑perfect deduction, high generalization for induction, and significant gains for abduction, while the design guarantees zero hallucination and supports continual learning.

By Weicai Huang (Beijing MQPat Technologies, Co., Ltd.)
arXiv AI
Aug 25

Measuring in-context algorithmic reasoning in language models against an exact Bayes-optimal reference

The paper introduces F-ICL, a benchmark that measures in‑context algorithmic reasoning in language models by exhaustively enumerating 86 million valid programs of length ≤13 on a Turing‑complete machine and computing the exact posterior under a bounded Levin–Solomonoff prior. Unlike typical benchmarks, F‑ICL provides a distributional reference rather than just answers, allowing the evaluation of models’ inductive priors. Across 105 configurations of models ranging from 0.8 B to 675 B parameters, models achieve up to 92 % accuracy, yet many still deviate from the Bayes‑optimal reference, and the study derives theoretical bounds on cumulative loss for predictors with positive prior weight on the reference.

By Luan Ozelim, Hector Zenil
arXiv AI
4d ago

How Much Prompt Is Enough? A Blackbox Minimization of Few-Shots in LLMs

The paper introduces ramework, a blackbox prompt‑minimization framework that identifies the minimal subset of few‑shot prompts necessary for large language models (LLMs). In a case study, the framework reduces few‑shot exemplars by an average of 65.3% in character count while maintaining full propositional output fidelity, revealing that models tend to keep logical identifiers and constraint declarations while discarding natural language prose. The analysis further distinguishes between universal encoder and decoder models, offering insights into prompt compression and structural analysis.

By Ali Alfageeh, Rahul Gopinath, Amin Alipour
arXiv Computation and Language
Sep 16

Autoformalizing Argumentative Material Inferences

The paper introduces GUARD, a neuro‑symbolic system that autoformalizes argumentative material by completing missing premises (guards) before formal verification. It uses large language models to generate candidate guards, Isabelle/HOL to verify them, and a contrastive test to ensure the proof depends on the original premises and does not over‑generalize. Experiments on Debatepedia and ARCT show that GUARD improves verified‑faithful scores by over 30 points and reduces leakage by about 20 points compared to prior LLM‑driven theorem proving methods.

By Xin Quan, Reto Gubelmann, Andr\'e Freitas