arXiv:2605. 11644v3 Announce Type: replace-cross Abstract: Positive data can show that two tuple occurrences share a successful sentence context without certifying that they are safely interchangeable.
By Takayuki Kuriyama
arXiv:2606. 07623v1 Announce Type: new Abstract: This paper develops a model-theoretic framework for verifying context-conditioned language-model behavior by replacing benchmark labels with finite semantic certificates.
By Faruk Alpay, Hamdi Alakkad
arXiv:2607. 06407v1 Announce Type: new Abstract: The XAI community has studied a wide range of queries and scores for explaining predictions of ML models.
By Marcelo Arenas, Pablo Barcel\'o, Diego Bustamante, Jose Caraball, Mar\'ia Alejandra Schild, Bernardo Subercaseaux
arXiv:2607. 21183v1 Announce Type: cross Abstract: The propositional abduction problem is a well-known form of non-monotonic reasoning where we are asked to find an explanation of a given manifestation.
By Johannes Schmidt (J\"onk\"oping University), Mohamed Maizia (J\"onk\"oping University, Link\"oping University), Victor Lagerkvist (Link\"oping University), Johannes K. Fichte (Link\"oping University)
arXiv:2604. 17621v2 Announce Type: replace Abstract: Many real-world questions appear deceptively simple yet implicitly demand two capabilities: (i) systematic coverage of a bounded knowledge universe and (ii) compositional set-based reasoning over that universe, a phenomenon we term "the tip of the iceberg.
By Xiao Zhang, Qianru Meng, Yongjian Chen, Yumeng Wang, Johan Bos
The paper introduces SAGE, a framework designed to reduce long‑horizon reasoning biases in large language models. It identifies two key biases—exploration bias and compounding bias—arising from complex reasoning spaces and sparse rewards, and proposes Symbolic Closure Analysis (SCA) to understand these effects. SAGE applies algebraic sparsification and hyperbolic structural guidance to suppress spurious branching and provide dense depth‑wise signals, achieving up to an eight‑fold improvement on the Andrews‑Curtis problem across multiple benchmarks and model families.
By Xinyue Zeng, Jiawei Zhang, Yujun Yan, Dawei Zhou